🥭 The Case of the Missing Mangoes: Testing Absence in Mumbai

:microscope:CUBE Chatshaala - Discussion Summary

Today’s session, held on the 4th of July 2026, brought together a familiar mix of curiosity and rigour, with Manali Bhujade, Aarya Takke, Arunan MC, Kiran, Prithviraj, and Sailekshmi among those present. The conversation moved between two very different biological threads, tied together by a shared underlying theme: what does it actually take to say something scientifically meaningful about the natural world?



The first thread began with Manali’s hypothesis around mango trees near Churchgate, Mumbai, specifically the claim that no mango fruiting would be observed in the area during July 2026. What made this interesting wasn’t the observation itself but the way it was framed and tested. Rather than simply walking past a mango tree and noting the absence of fruit, the group discussed the importance of drawing a representative sample. Five mango trees were selected from on and off Pedder Road, treated as a stand-in for the broader population of mango trees in that part of the city, in order to give the hypothesis a fair chance of being properly evaluated rather than resting on a single anecdotal glance.

Alongside this, the discussion referenced a related whiteboard map plotting locations across Mumbai, including Malad, Thane, Vaibhavwadi, Colaba, Marine Lines, and Elphinstone College, situating the mango-fruiting question within the wider geography of where such observations were being made and compared.

The second thread of the day concerned germination biology. A batch of 100 green gram (moong) seeds was soaked in water for 24 hours, and at the end of that period, 90 of the 100 seeds had sprouted, giving a 90 percent germination rate within the first day. The sprouting data was also broken down across four separate trials or batches, showing considerable variation: 8 out of 10, 7 out of 10, 10 out of 10, and 5 out of 10. This spread naturally opened up a discussion on why germination rates fluctuate even under seemingly identical conditions, touching on seed viability, moisture uptake consistency, and the limitations of small sample sizes when trying to draw general conclusions.

Underpinning both experiments was a conversation about the hypothetico-deductive (HD) method itself, the classical framework in which a hypothesis is proposed in advance of data collection, based on existing theoretical understanding, and is then subjected to deliberate attempts at falsification rather than simple confirmation. As one recent review in PMC describes it, a well-constructed hypothesis is central to scientific knowledge, guiding the research process from a problem toward its potential solution, with the hypothetico-deductive framework serving as the critical link between theory and empirical testing. The same review points to a useful checklist for hypothesis quality, sometimes called the 5E rule, which frames an effective research hypothesis as Explicit, Evidence-based, Ex-ante, Explanatory, and Empirically testable. This framing gave the group a useful lens for revisiting Manali’s mango hypothesis: was it explicit enough? Was it truly ex-ante, formulated before observation rather than fitted to it afterward? And crucially, was it falsifiable, meaning was there a genuine possibility that the data could have proven it wrong?

This last point matters a great deal in citizen science. A hypothesis that cannot fail is not really scientific at all; it is closer to a description. The mango and moong experiments, in their own modest ways, illustrated what it looks like to build claims that could have gone either way, and then to check.


:red_question_mark: Provocative Questions

  1. If the sample of five mango trees had shown fruiting on even one tree, would that have been enough to falsify the hypothesis outright, or would the group have questioned whether that tree was representative?

  2. Why might germination rates vary so much between batches (8/10, 7/10, 10/10, 5/10) when the seeds were sourced and soaked under what appears to be the same protocol? What hidden variables might explain that spread?

  3. Is “representative sampling” from five trees along and off Pedder Road genuinely representative of mango trees across Mumbai, or does it only tell us something about that specific stretch of road?

  4. How would the mango-fruiting hypothesis need to be reworded to satisfy all five criteria of the 5E rule: Explicit, Evidence-based, Ex-ante, Explanatory, and Empirically testable?

  5. What would it take to turn the moong seed germination observation from a one-off data point into a properly falsifiable hypothesis with predictive power for future batches?

  6. Given that no mango fruiting was hypothesised for the whole of Mumbai in 2026, what would a single counterexample actually prove, and would it be enough to overturn the claim?


:black_nib:What I Have Learned

This session was a helpful reminder that formulating a hypothesis is a skill in its own right, separate from the excitement of collecting data. It’s tempting to treat any observation- a lack of mango fruit, a pile of sprouted seeds- as self-evidently meaningful. But the real discipline lies upstream of that, in stating clearly and in advance what we expect to find and under what conditions we would consider ourselves wrong.

I found the 5E framework genuinely clarifying. It gives a concrete way to interrogate a hypothesis before any data collection begins, rather than after the fact when it’s easy to retrofit a story onto whatever numbers happen to show up. The mango hypothesis, on the surface a simple claim about fruiting season, turned out to be a good test case for thinking about sample representativeness, geographic scope, and the timing of the claim relative to the observation.

The moong seed data was a good complement to that discussion. A 90 percent germination rate sounds tidy and conclusive until you look at the batch-level breakdown and see numbers ranging from 50 to 100 percent. That variation is the real story, and it’s a useful lesson in not letting an aggregate figure smooth over the noise underneath it.


:glowing_star:TINKE Moments (This I Never Knew Earlier)

The clearest TINKE moment of the session centred on the distinction between an observation and a hypothesis. Simply noting that a mango tree has no fruit is an observation. Stating in advance that no mango fruiting is expected across Mumbai in July 2026, and specifying the conditions under which that claim would be considered false, is a hypothesis in the proper HD sense. The group’s discussion made explicit something that had perhaps been implicit before: a claim only counts as scientific once there is a real chance it could be shown wrong.

A second TINKE moment arose from the batch-wise germination data. Before this session, it might have been tempting to report only the combined 90/100 figure and treat that as the full picture. Breaking the results down by batch revealed that variability, ranging from 50 to 100 percent sprouting, is itself informative and shouldn’t be discarded in favour of a single tidy average.

A third moment came from the sampling discussion around the mango trees. Choosing five trees “on and off Pedder Road” as representative of a much larger population is a judgement call, and today’s conversation made explicit the reasoning that has to sit behind that kind of choice, rather than treating sample selection as an afterthought.


:warning: Gaps and Misconceptions

One gap that surfaced was around what “representative” actually means in a sampling context. A sample size of five trees from a single road is a reasonable starting point for an informal citizen science observation, but the discussion didn’t fully resolve how such a small, geographically narrow sample should be interpreted when the hypothesis makes a claim about mango fruiting across the whole of Mumbai. There’s a distinction between a sample that is representative of a local neighbourhood and one that can support a city-wide claim, and this distinction deserves more attention in future sessions.

A related gap concerns the moong seed protocol. The whiteboard notes record the outcome, 90 out of 100 seeds sprouted after 24 hours of soaking, along with four batch-level breakdowns, but the underlying conditions (seed source, water temperature, container type, whether batches were run simultaneously or sequentially) were not fully detailed. Without that information, it’s hard to know whether the variation between batches reflects genuine biological variability or differences in experimental handling.

Finally, there’s a subtle misconception worth flagging: treating a null hypothesis, such as “no mango fruiting is expected,” as though it requires less rigour than a positive claim. In the hypothetico-deductive framework, a claim of absence is just as falsifiable, and just as much in need of clear operational definitions, as a claim of presence. The session’s engagement with the 5E rule was a useful corrective here, though it’s a point worth returning to and reinforcing in future discussions.


:camera_with_flash: Photographs during Chatshaala

:books: Reference