Results From Our Farm Program Outcome Evaluation
- Jennifer-Justine Kirsch

- 2 days ago
- 6 min read
Highlights
Process: We ran a randomized controlled trial with 53 farms to evaluate our farm program’s outcome: water quality improvements.
Goal: Determine whether our intervention, rather than external factors like weather, was responsible for those improvements.
Results: Treatment farms (which received their readings and tailored recommendations) saw water quality improve 82% of the time, compared to 17% of the time for control farms (which received neither readings nor recommendations). This improvement was measured on the third day after a water quality issue arose.
Conclusion: The ARA's core mechanism seems to work, though some questions remain open, including how much farmers' own, unprompted actions contributed.
View the full study report:
Background: Our Farm Program
The Alliance for Responsible Aquaculture (ARA) is FWI’s farm program. We work with fish farmers in India to monitor water quality and stocking densities, and provide corrective actions when conditions fall below ranges optimal for fish welfare. Once water quality improves, we consider whether the event should be counted towards our fishes helped estimate.

Motivation for This Study
Our welfare science research showed that maintaining water quality in certain ranges benefits fishes. What it could not tell us is whether our ARA intervention improves water quality to these ranges, or whether outside factors like weather would have produced the same result.
The outcome evaluation was designed to answer this question directly: Does providing farmers with water quality readings and corrective actions result in better water quality?
We ran it from February to May 2026, after an initial attempt the previous year failed due to protocol errors leading to poor data quality. Fifty-three farms participated: 25 in the treatment group (receiving regular ARA measurements and corrective actions) and 28 in the control group (receiving neither). We measured water quality in both groups every two weeks and followed up after 2 and 3 days whenever readings were out of range. The key metric was the "resolve rate": how often water quality had returned to acceptable welfare levels by Day 3.
For the full methodology, see the study protocol.

Study Findings
The ARA's water quality intervention works, with one question left open. When water quality fell out of range, treatment farms brought it back within welfare limits by Day 3 in 82% of cases. Control farms did so only 17% of the time—a 65-percentage-point gap.

Three separate analyses of the same dataset reached this conclusion. Two were run by external analysts (Adam Holt and Norma Forero Muñoz), who were blind to group assignment during the primary analysis to reduce confirmation bias. A third analysis was run by our study lead using Claude, an AI assistant (this analysis was not blinded). We also ran several statistical tests confirming the gap is very unlikely to be down to chance, and found no evidence that weather explained it.
For more details, see the full study report, including all the statistical analyses:
Limitations
The headline result is clear, but its scope is limited. Below are the main limitations to interpreting the results of this study:
Self-initiated actions: The main uncertainty is about how much of the water quality improvements resulted from farmers' own actions. Given the complexity and importance of this limitation, we discuss it further in the section below.
Duration: Water quality improvement was almost exclusively visible at the Day 3 follow-up; Day 2 resolution rates were similar between control and treatment farms. We do not have information on whether the resolution rate in both groups would have changed beyond Day 3.
Previous exposure: Every participating farm was an existing ARA member, since recruiting farmers new to the program was not feasible. The study could not tell us about the treatment effect on ARA-unexposed farmers. The majority of farms had also participated in last year’s iteration of this study and flipped in their control/treatment assignment.
Seasonality: The study ran for three months in summer. We do not know whether the effect holds across all seasons.
Geography: Farms came from a single area. Results may not generalize to other regions, or even other parts of the state where we operate.
Parameters: Most out-of-range readings involved dissolved oxygen. The findings tell us less about pH and ammonia, two other key welfare indicators.
How Much Did Farmers' Own Actions Contribute?
Unfortunately, the study data were not conclusive on the effect of actions that farmers took themselves. Farmers in both groups sometimes took steps to improve water quality without any prompting from us. We call these self-initiated actions (SIAs), and we tracked them because they were the most plausible alternative explanation for our results.
Only 56% of recorded SIAs came with a specific implementation date; for the rest, farmers could give only a rough window, which made it hard to line an action up with the improvement that followed. These were our main findings on SIAs:
SIA rates were fairly similar across groups overall, though treatment farms reported them somewhat more often specifically during out-of-range events (around half of out-of-range visits, versus about a third at control farms). This does not point clearly in either direction: it is consistent with SIAs mattering in both groups, and it does not separate their effect from our corrective actions.
In the control group, where farmers received no corrective actions, SIAs that targeted the out-of-range parameter did not resolve it on their own in any of the 10 events we recorded. That could mean SIAs alone rarely fix these problems, or it could mean some of those SIAs happened too late in our three-day follow-up window for us to detect an effect, since we could not date roughly half of them. We cannot tell which explanation is right from this data.
In the treatment group, every farm that experienced poor water quality received a corrective action. If they implemented this and their own SIA, we could not fully distinguish which one caused the treatment effect. While our fish welfare expert panel suggested that, in most cases, corrective actions are the major driver (see below), we do not have primary data for this.
Put together, this leans significantly toward corrective actions mattering more than SIAs, though we cannot rule out a meaningful SIA contribution. Either way, the core comparison still holds: treatment farms resolved out-of-range events far more often than control farms, measured over the same three-day window the ARA already uses in practice. Combined with our earlier research on these water quality ranges, this adds to the evidence that the ARA is improving fishes' lives. Our program team is now considering a follow-up study to isolate and compare the effects of corrective actions and SIAs.
Lessons from Running a Field RCT
Because this was our second attempt at this study, we came away with several lessons we hope are useful to others doing similar work:
Take time for protocol development: We had to repeat this study because we realized two major issues with our protocol that affected the quality of data we received. Looking back, we should have spent more time stress-testing the protocol, for example, through piloting it (see point below).
Pilot longer than you think you need to: Our two-week pilot was not enough. A one-month pilot with a thorough data review at the end would give more time to catch protocol problems before they affect the main study.
Build in strong quality controls: Back-checks to confirm staff are implementing tasks correctly were essential for study integrity. We copied most of these from regular ARA programming and found it useful to incorporate these existing, tested protocols into the study.
Implications for the ARA’s Future
This study gives us more confidence that the ARA's core mechanism works, which is a meaningful step forward. But scaling the program remains a challenge.
The ARA currently reaches 11 fishes helped per dollar, below our internal cost-effectiveness threshold of 20. We have not found a scalable delivery mechanism that would change this: our satellite imagery project was not successful, and we are still relying on in-person farm visits, which are hard to scale cost-effectively.
To reach our 2026 goal, we are pursuing improvements to the ARA (including a water quality prediction model currently in testing) and running separate programs that may offer a better cost-effectiveness profile, including feed fortification and chill-kill slaughter. We will share updates on these on our blog in the coming months.




Comments