Ai2 & Asta: Letting AI Surprise Itself
Deep Dive | Edition 24
Welcome back to the deep dive, where we break down the AI tools and data reshaping how new drugs are discovered. In each edition, we speak directly with the teams behind these tools to explain what they solve, how they work and where they are going next.
What happens when you stop asking AI to confirm your hypotheses and start asking it to find its own?
For: Computational biologists, genomics researchers, data-heavy lab groups, and anyone sitting on large datasets they suspect contain more than their pre-registered hypotheses will find.
Ai2’s AutoDiscovery tool explores datasets autonomously, scoring hypotheses by how much they shift an AI observer’s beliefs, not just statistical significance.
It matters because the bottleneck in data-rich fields is no longer compute or analysis speed. It’s ideation: deciding what questions to ask in the first place.
We spoke with Bodhi Majumder, Senior Research Scientist at Ai2, about searching for surprise and what it means to let a model find things it didn’t expect.
You read this newsletter because staying current in AI x life science matters to you, it matters to us too.
That’s why we’re building a free personalised resource for the community: from papers, patents, and clinical trials to c-suite career moves and job opportunities. So you can stay up to date with what actually matters to you.
We want to know: where else are you looking? What’s missing? Two minutes of your time will shape what gets built.
Today we spoke with Bodhi Majumder at Ai2 about Asta, their ecosystem of scientific tools, and specifically AutoDiscovery, a system (now available as an experimental feature in AstaLabs) that explores datasets autonomously to find hypotheses that even the AI didn’t expect.
The problem: datasets are growing faster than the questions we know to ask
Datasets are getting bigger. Cancer genomics, brain activity, population health surveys. The way scientists explore them hasn’t really kept pace. You come in with pre-registered hypotheses, test them, and publish. Whatever falls outside your initial framing stays buried.
The issue isn’t compute. It’s focus. If you’re fishing for salmon, you’ll find salmon. You won’t notice the unexpected species swimming right behind it.
“It’s not about time,” Majumder explains. “It’s about the frame of reference. Maybe you’re just not thinking from a particular angle that allows you to find something completely different.”
So the bottleneck is ideation. Models are getting very good at executing tasks you hand them, but deciding what to work on in the first place still requires human creativity, and human creativity is bounded by what you already know to look for.
The idea: optimise for surprise, not significance
AutoDiscovery takes a hypothesis-free approach. You upload a dataset, optionally describe the schema and context, set a budget (say 100 hypotheses), and let it run overnight. Each hypothesis takes about four minutes.


The metric it’s optimising for is Bayesian surprise. The system uses a language model as the reference knowledge frontier. For each hypothesis, it first asks: given your knowledge of the world, what’s the probability this is true? Then it goes into the data, writes code, runs statistical tests, and asks the same question. If the final belief shifts dramatically (say from 70% to 20%), that counts as a surprise.
“We’ve found that 65% of insights the AI finds surprising are also rated as surprising by human experts,” Majumder says.
That 65% number is worth pausing on. It means the system isn’t just flagging noise or artefacts. Nearly two-thirds of the time, what surprises the model also surprises the domain expert. The remaining third probably represents cases where the model lacks context the expert has (which is itself a useful signal about where the model’s world knowledge breaks down).
What it’s finding
The tool is domain-agnostic. Researchers have now used it across oncology, marine ecology, and social science.
At Providence Swedish Cancer Institute, Dr. Kelly Paulson’s team ran it on breast cancer and melanoma datasets. It confirmed what they expected (immune activity matters in melanoma, PI3K pathway in breast cancer) and then surfaced associations they hadn’t been looking for, including potential links to lymph node metastasis risk in breast cancer. Those novel hypotheses are now in follow-up validation studies. That’s the ideal outcome: the system finds something you wouldn’t have thought to ask about, and it’s specific enough to actually test.
At Scripps, a marine ecology group used 20+ years of rocky reef monitoring data from the Gulf of California. Auto-Discovery identified relationships between productivity across trophic levels that would have required extensive manual work to find. In social science, a University of Utah researcher used it to discover that doctoral-degree holders edited AI-generated writing substantially more than those with undergraduate or master’s degrees. That one is already peer-reviewed and published (November 2025).
“It not only helps with efficiency,” Majumder says. “It helps with coverage. It discovers unexplored ideas that would have taken a very long time for a human to find.”
The Asta ecosystem
Auto-Discovery is one piece of Asta, AI2’s broader platform. Since our conversation with Bodhi, the workflow between the two has become more direct: researchers can now select “Explore with Asta” on any Auto-Discovery hypothesis and it hands off directly to DataVoyager (Asta’s agent for data-driven discovery and analysis) with the dataset and results already loaded. No manual transfer.

“That’s one of my favourite workflows,” Majumder says. “Auto-Discovery finds something interesting autonomously. Then I load it into Asta and have a conversation about it.”
What’s nice about this (and I think somewhat underrated as a design choice) is that the scientist gets to choose their window of autonomy. Auto-Discovery runs unsupervised. Asta’s chat mode is collaborative. You decide how much rope to give it.
Asta’s literature search also recently added a “Deep search” mode that evaluates whether results actually answer your query and keeps searching if they fall short. Auto-Discovery itself is now available as an experimental feature within AstaLabs.
Why it’s different
Most AI tools for data analysis start with a question. This one starts with a dataset and tries to find the questions worth asking. Optimising for Bayesian surprise rather than p-values alone means it surfaces things that challenge your existing assumptions, not just things that clear a statistical threshold.
It’s also publicly available (https://autodiscovery.allen.ai/runs, 500 free credits, no coding required) and the code is open source for researchers who want to mess with the search objectives themselves.
The future
Majumder sees the biggest opportunity in closing the ideation loop. The team is exploring live literature integration and expanding beyond surprise to other search objectives (utility-driven discovery, confirmatory hypotheses). The longer-term direction is semi-autonomous workflows where scientists tune how much autonomy the system gets and at what stage.
“We feel the biggest value comes from semi-autonomous loops where the scientists decide the window of autonomy,” Majumder says. “Human-AI co-creation of ideas. Not replacement of scientists.”
Kiin’s view
“Optimise for surprise” is a clean, testable objective, and it sidesteps the usual problem with AI-for-science tools (that they only find what you already suspected, just faster). The 65% agreement rate with human experts is encouraging. The real question is whether surprises translate into publishable, mechanistically validated findings, or whether they mostly end up as interesting statistical anomalies that don’t lead anywhere. The domain-agnostic claim is ambitious. We’ll be watching the oncology and neuroscience partnerships to see if they produce papers that actually shift thinking.
🧑🔬 Get in touch with Bodhi.
💻 Ai2 Website.
🔍 Try AutoDiscovery.
🧪 Explore Asta.
Thanks for reading Kiin Bio Weekly!
💬 Get involved
We’re always looking to grow our community. If you’d like to get involved, contribute ideas or share something you’re building, fill out this form or reach out to me directly.
Subscribe now to stay at the forefront of AI in Life Science and keep up with this upcoming season of deep dives.
Connect With Us
Have questions on this or suggestions for our next deep dive? We’d love to hear from you!
📧 Email Us | 📲 Follow on LinkedIn | 🌐 Visit Our Website



