Mass Dynamics Blog

Beyond chatbots for science: your agent is only as good as its room

Written by A/Prof Andrew Webb, Ph.D | July 29, 2026 at 2:23 AM

The capability is not the variable. The room it works in is - shared data, visible methods, and every claim carrying its warrant with it.

 

The most capable tool science has ever been handed is arriving faster than our ability to know when and how to trust it. Used carelessly, it can produce plausible error at a scale and speed our existing safeguards were never designed for. But that is not the whole story. Harnessed correctly, AI could reason across more evidence than any scientist or team can hold, expose assumptions we routinely leave unstated, and strengthen rather than weaken the foundations of scientific trust. Learning how to build those systems will take years of trial, error and adaptation, even as the capability itself continues to accelerate. The challenge is not simply to make science faster. It is to ensure that the speed of science does not outrun its warrant.

 

A team was running large-scale dose-response curves across a compound screen, chasing newly identified targets from a phenotypic screen. One series of compounds was producing unusually stark shifts in biology. And the person who spoke up was not the analyst. It was a chemist, who asked, almost in passing, whether the targets shifting in the assay were kinases.

 

That is the kind of question that normally goes into a queue. A bioinformatician takes it away, and the answer comes back days or a week later, by which point the meeting that prompted it is long over and the thread of intuition that produced it has gone cold. Instead, because the data was live and shared in the room, the question was answered while it was still being asked: a handful of clicks, and yes, they were kinases. That single resolved question changed the direction of the conversation in the moment, and with it the trajectory of the screening campaign, in a way that plausibly saved weeks and potentially millions of dollars.

 

The team did not wait three weeks to have the real conversation. They had it while the question was still warm, and they had it with the caveats in the room rather than discovered later.

 

This is roughly what I think the near-term future of science looks like. Not an autonomous "AI scientist" that vanishes into a black box and comes back with the answer. Not a chatbot handing over a polished interpretation. Something closer to a room: a shared environment where the right people, data, methods, context and AI agents can reason together, and the evidence behind a claim stays in view.

 

At ASMS (the biggest mass spectrometry conference in the world) this year, AI was a central topic in nearly all my conversations and catch-ups. I heard it referenced in the talks, the posters, the booths, the hallway conversations. But when I asked people what they were actually using it for, the answers were strikingly small relative to what AI agents and harnesses can do right now. Cleaning up an email. Drafting a bit of code. Summarizing a paper they hadn't had time to read. Useful things, but careful, peripheral ones. None of the scientists I spoke to were willing to let AI near the part of the work that matters most to them: the interpretation.

 

That gap, between how present and capable AI suddenly is and how little we yet trust it with, is what this post is about.

 

From answers to systems that take part in the work

Most scientists are using AI much as everyone else does: a chat window, a question, an answer. Sometimes the answer is useful. Sometimes it is wrong. Often it is fluent enough that the first instinct is not to ask whether it is useful, but whether it can be trusted at all. This is where most scientists are starting, not because they lack imagination, but because it is how AI has been packaged. It is the most available interface. It is easy to try. It is also easy to outgrow.

 

A smaller group has started wiring AI into the actual machinery of research: literature synthesis, analysis workflows, experimental planning, interpretation - increasingly with agentic systems that can loop, check their own work, hold context and act across tools.

A chatbot answers a question and forgets you the moment the window closes. An agent is slightly more advanced: it can act across tools, check its own work, and - if you let it - hold on to the context of who you are and what you are working on. The interesting frontier is not a smarter one-shot answer. It is an agent that stays, learns the relevant, specific lab context rather than lab-in-general, and carries that understanding forward from one question to the next.

 

This is not a fringe activity, and it is no longer only the technology companies saying so. Earlier this year, Nature's editors gave a leading editorial to the arrival of "AI scientists", arguing that institutions, funders and publishers will have to rethink how research is organized, credited and governed. But it is not shaping up the way the marketing brochures suggested. What I see coming is an augmented future, where scientists spend less time in the weeds and more time on the concepts, the scientific ledger, and the iteration and honing of new ideas.

 

Once a model sits inside that kind of harness, it stops behaving like a conversational interface and starts behaving like something that can reason its way across a problem. So the question is no longer whether scientists will use AI. They already are. The question is what we are asking it to do. What class of scientific work are we inviting AI into?

 

Cleaning up an email is one thing. Writing a helper script is another. Interpreting complex biology is something else entirely. In my last blog post I argued that AI for science will not unfold like AI for code: the feedback loops are slower, the ground truth stranger, the context more tacit, the standards of trust higher. Code can be tested against a specification. Science has to reason against a universe we are still discovering.

 

Here, I want to go one layer deeper. Not the peripheral uses - the code, the summaries, the admin - real as their value is, but something more delicate and more consequential: AI for scientific interpretation.

 

By scientific interpretation, I mean the work of moving from complex experimental data to biological meaning. From hundreds of changing proteins or post-translational modifications to a plausible mechanism. From a noisy multi-omic dataset to a decision about what to test next. From a pattern in the data to a claim another scientist, clinician, reviewer, investor, or regulator might be willing to act or build upon. That is a very different class of work, and if we blur the distinction we risk importing the wrong assumptions about speed, automation, and trust into the most judgment-heavy layer of science.

 

This is where I think we need both caution and ambition. Cautious, because a fluent model can be wrong. A confident interpretation can collapse under the next experiment. Biology does not care how elegant the story sounds. But ambitious, because the frontier of AI is no longer just a better autocomplete. We are beginning to see systems that can search broadly, hold a large amount of context, weigh competing explanations, and coordinate work across tools. As in software engineering and cybersecurity, put a capable model inside the right loops and harnesses and it can reason across large, messy systems in ways that genuinely transform what is possible.

 

And here is the thing that decides whether any of this capability is worth anything: the same agent, on the same task, is either careful or useless depending almost entirely on what it can see. Give it the raw file and nothing else and you get confident nonsense. Give it the dataset, the methods, the caveats the lab already knows, the history of what has been tried - and the same model reasons like a careful colleague. The capability is not the variable. The room it works in is. That is what I mean by the title: your agent is only as good as its room.

The bottleneck is interpretation

At the edge of biology, the limiting factor is rarely a single missing fact. It is the diffuse, nuanced, contested process of connecting observations, mechanisms, prior knowledge, experimental caveats, and competing interpretations. It is a place where language matters, context matters, and inference is hard.

 

That is precisely why AI could become so powerful for interpretation. Not because it replaces the scientist, or magically produces truth, but because a well-designed human-AI system may be able to reason across more evidence, more context and more prior work than any scientist or small team could hold in their heads at once.

 

When you allow yourself to step back far enough, all of science becomes one long act of conversion: turning observation into understanding, and understanding into things that matter, medicines, harvests, materials, diagnoses that arrive in time. And the rate of that conversion has always been bounded by the same narrow thing: how much a human mind, or a small room of them, can hold at once. Every discovery ever made had to fit through that aperture. It is biologically fixed, and in four hundred years it has not budged.

 

Today, the generation of data is under no such constraint. In my own experience, a single well-run experiment now produces more measurements than an entire career could once have examined by hand a few decades ago. The bottleneck of science is no longer the data. It is interpretation: the slow, human work of deciding what the data means. And that bottleneck is not widening to meet the flood. It is roughly fixed, with some pretty hard limits, while everything upstream of it accelerates.

 

Which makes something quietly staggering almost certainly already true. Somewhere, in most large datasets scientists have already generated, sit undiscovered patterns that could change lives. A mechanism. A target. An answer to a disease. Not waiting to be discovered so much as already discovered, in the only sense that matters, the evidence exists, and simply never noticed, because no one had the attention, awareness or breadth of comprehension to see it.

 

The real prize, then, is not performing today's science faster. It is changing the rate at which our species turns what it has measured into what it understands - unpinning that rate, for the first time, from the bandwidth of a single human brain.

 

The technology itself is completely neutral on which way this goes. The same systems that could change that conversion rate can also flood science with plausible, confident noise, superficial certainty traveling faster than the evidence underpinning it. What separates the two outcomes is not the LLMs. It is how much rigor we build into the systems that carry them - how much scrutiny, provenance, and visible method travels with every claim.

 

Build it loose, and AI becomes a machine for manufacturing slop faster than we can check it. Build it with rigor, and interpretation stops being a series of isolated expert acts that evaporate when the meeting ends. Ideas begin to compound, accumulating across people, datasets and years, the way capital compounds and the way knowledge was always supposed to, but rarely does.

Science is right to be cautious

Scientists can be skeptical about new technology. In my experience, sometimes painfully so. But skepticism is not a quirk of scientific personalities. It is the discipline's protective immune system. And AI is the first technology I've encountered that can potentially slip past it.

 

The apparatus of the field - the controls, the replication, the peer review, the demand that a result hold up in someone else's hands - exists for one purpose: to keep us from being fooled, including by ourselves. A scientist's first instinct in front of a confident claim is to ask how it could be wrong. That instinct is not a flaw to be coached out of them. It is the thing that makes science trustworthy.

 

The immune metaphor is not an accident for me. I spent more than a decade among some of the world's leading immunologists, studying how a system learns to tell self from non-self, signal from noise, and how finely tuned it has to be: protecting us from life-threatening infection one moment, capable of turning on the body the next as autoimmunity, or tipping into the kind of catastrophic overreaction I spent my post-doc at Imperial studying the immunopathology of hemorrhagic dengue fever.

 

The same machinery, producing protection or harm depending on regulation and context. That is the part that stayed with me, and it is the part I keep seeing again now. What makes an immune system dangerous is not that it is weak. It is that it is powerful and can be badly regulated. Its capability was never the question - the context it operates in decides whether that capability defends you or destroys you.

 

That training left me with something else, too: how I read my own results. The most dangerous error is rarely the one someone else catches. It is the elegant one you fall in love with, the result you want so badly to be true that you quietly stop asking how it could be wrong. The skepticism I am describing is not abstract to me. It is the habit that survived.

 

Which is exactly why AI unsettles the best scientists. The danger was never that a model would be wrong; science is built to catch wrong. The immune system evolved to catch pathogens that announce themselves - the friction, the effort, the tells of something that had to fight its way into being. AI produces a claim with none of those signatures: clean, plausible, confident, and wrong. It makes the surface of a claim independent of the work beneath it - precisely the thing the immune system cannot see easily. And there is something almost evolutionary in this. In training models to give answers people rate highly, we have inadvertently selected for outputs that do not trigger a scientist's skepticism - fluent, confident, well-formed - regardless of whether the substance beneath earns that confidence. Not a system that fools us on purpose, but one shaped, like an organism adapting to its host, to pass the checks we happen to run at the door.

 

So when a scientist recoils from a smooth answer that arrived too easily, they are not failing to understand the technology. Their immune system is working correctly, and the discomfort is telling them something true: the claim has not yet earned belief.

The answer to a trained skeptic is never "trust me." It is "here is what you can check."

There are two kinds of skepticism, and only one I believe will be useful. One rejects the new tool on sight, as a threat to rigor. The other asks the harder question: what would this tool have to show me before I believed it? The first protects nothing; the second is how we might advance science beyond that human bottleneck.

The old workflow was static

The traditional workflow for complex scientific data is often a chain of translation.

The instrument produces files. A specialist processes them. A spreadsheet appears. Figures are made. A slide deck is assembled. A meeting happens. Questions arise. The specialist goes away to reanalyze. Another meeting is scheduled. We loop until we move onto the next project.

 

This can work; it has done for decades. But it creates latency at exactly the wrong place: the moment when curiosity is highest and the team is most ready to reason together. By the time the answer comes back, the chemistry has moved on. Biology has moved on. Sometimes the decision has already been made.

 

We see this again and again with customers. Once complex omics data moves out of static spreadsheets and into a shared environment, the conversation changes: people stop asking for summaries and start asking better questions of the data itself.

 

None of this means faster is automatically better. Some friction is load-bearing: a room that collapses the distance between question and answer too far can manufacture false consensus, a team agreeing quickly because the data appeared quickly, not because it earned the agreement. So the goal is never simply to remove latency. It is to remove the bad latency - the reformatting, the waiting on a specialist, the re-runs, the translation between disconnected tools - while keeping, and ideally sharpening, the scrutiny.

That is not just faster analysis. It is a different shape of collaboration.

Inside the room

Science often looks solitary from the outside. A person at a bench. A person at a computer. A person writing alone late at night. But trustworthy science is rarely produced alone. It is produced through the friction of competing minds, through disagreement, through the slow process of exposing an idea to people who know different things, notice different weaknesses, and carry different histories of what has failed before.

 

I believe AI can help with this too, and in time may help enormously. But only if it is working in the right environment.

 

The most useful thing I have learned from working with AI agents every day, the coding ones included, is almost embarrassingly simple: context is everything. And science runs on a vast amount of context that never gets written down. We leave it between the lines, shorthand that is efficient between experts with a shared history, and invisible to anyone, or anything, that does not.

 

So when I talk about a room, I am really talking about a place where that context becomes legible: where the assumptions underneath a claim are captured rather than implied, and the evidence behind an interpretation can be inspected rather than taken on trust. There is more latent untrustworthiness in science than we care to admit - not through bad faith, but because errors slip through when too much is left unstated. As we generate more data than any of us can hold in our heads, leaving things between the lines stops being efficient and starts being risky.

 

That environment is the room. Not necessarily a physical one, and not a new platform to log into either. Increasingly it can live inside the tools a team already works in, so the room comes to them rather than the other way around. What makes it a room rather than a folder of files is that everything the science depends on is present and live in the same place, at the same time.

 

  1. The data stays live. Not a spreadsheet on one person's laptop, exported into static figures that are already stale by the meeting, but the actual data, open in front of everyone, still queryable while the conversation is happening. When someone asks "what if we drop the flagged samples?", the figure changes in the room, not in a follow-up email three days later.
  2. The methods stay visible. Not buried in a script only the analyst can read, but legible enough that a biologist can see which normalization was used and a chemist can ask why - and get an answer in the moment, not in a queue.
  3. The context persists. Every lab carries context that does not live cleanly in a file: the run everyone knows not to trust, the reagent batch that behaved strangely, the normalization choice that makes sense only if you understand the sample prep, the pathway list that matters because of a failed experiment three years ago, the caveat that never makes it into the methods section because everyone in the group "just knows". Today that knowledge lives in people's heads and walks out of the door when they do. The version of this I find most compelling is not a room that stores it passively, but an agent that holds it - something that remembers the flagged batch, the failed experiment, the correction someone made last quarter, so the lab does not have to carry all of it in living memory. Interpretation stops resetting to zero every time a person moves on.
  4. The interpretation stays open, where it can be challenged. A neuroscientist enters a dataset through one set of proteins, a chemist through target engagement, a clinician through phenotype, a computational scientist through variance and effect size. Progress usually depends not on eliminating those differences but on making them visible enough to argue with. So the most useful system is not the one that returns a single polished paragraph. It is the one that lays out the competing interpretations, shows the evidence for each, and helps the team decide what to test next - sharpening the debate rather than bypassing it.
  5. AI has a seat in the room too, but not as an oracle handing down answers. More like a set of capable colleagues with the same context and the same boundaries - searching the literature while the question is still warm, checking whether this pattern showed up in last year's experiment, flagging that the headline result flips if you move the threshold, asking, quietly, whether the interpretation on the table is defensible. The best of these systems will not be the ones that generate the most impressive answer, but the ones that work inside the standards of a specific team, method and dataset, and flag when a choice departs from precedent.
  6. And the decision leaves a trail. The next scientist does not begin from a static PDF or a half-remembered meeting, but from a living environment that already holds the data, the methods, the interpretation and the disagreement that shaped the last decision. Interpretation becomes cumulative rather than disposable - which, in a field generating data faster than anyone can read it, is the only way it keeps pace. A chatbot gives you an answer and the answer leaves with you. A room helps a group build understanding that stays - and, built well, understanding that compounds, each question answered making the next one more valuable.

Warrant is built in the open

I've begun to refer to warrant as the auditable reason a stranger should believe a scientific claim. I keep using it because I think it is one of the most important ideas in this new age where we are co-existing with AI.

 

We used to set the bar for warrant ourselves. AI raises it.

 

When a model can generate plausible scientific prose in seconds, fluency stops being evidence. Confidence stops being evidence. Even coherence stops being enough. The question becomes: what can the system show? Can it show the data behind the claim? Can it show which method was used, and why? Can it show whether another method would have changed the conclusion? Can it show which assumptions were made, what the human accepted, rejected, or overrode, and where the interpretation actually came from?

 

A chatbot is a poor container for that kind of work. A room can hold all of it - everything above, plus the provenance and the decision trail - and let a claim carry more of its history with it. This is the layer beneath the room I'll come back to: the substrate that makes an environment scientific rather than merely shared. The future of AI in scientific interpretation should not be about generating more claims, but claims that are easier to inspect, easier to challenge, and easier to trust.

The scientist remains central

The old worry - that AI would replace the scientist, dilute expertise, or flatten hard-won judgment into generic output - has quietened, and rightly so. What's replacing it is a more interesting uncertainty: not whether there will be a role, but what the role becomes. The closest guide I have is software engineering. As the tools got dramatically more capable, demand for good engineers didn't collapse, it grew, because the scarce thing was never typing code. It was the systems thinking: knowing what to build, how the pieces fit, where it breaks, what to trust. The manual craft got absorbed; the systems level judgment became more valuable.

 

I think science will follow the same path. The most powerful version of AI here does not remove the scientist; it makes judgment the binding constraint - the thing that guides, constrains, challenges and signs off on the work of increasingly capable systems. The scientist becomes the person who knows what question is worth asking, when the data is good enough to act on, which caveats are fatal and which are manageable, and what evidence would actually change the next step.

 

And once a scientist has made that call, the room need not stop at understanding. The same environment that helped interpret the data can help carry the decision into the next step - running the approved workflow, drafting the report, opening the follow-up - with the human signing off and a full record of what was done and why. The point is not that the agent acts on its own. It is that nothing is lost in the handoff between deciding and doing, and the scientist stays the one who decides.

 

The future I'm excited about is not one where AI has more agency and scientists have less. It is one where scientists are elevated to spend more time thinking, where AI absorbs the friction around them: the searching, reformatting, cross-checking, rerunning, comparing, annotating and remembering. The scientist is not reduced. The scientist unleashed.

Scientific Intelligence is a system, not a model

At Mass Dynamics, we have started using the phrase Scientific Intelligence for this broader idea. Not intelligence as an algorithm. Intelligence as a system: people, process, data, methods, infrastructure and AI working together so scientists can move from complex, multivariate data to confident, defensible decisions. This distinction is important because the AI conversation is still too model-centric. Better models will help. Of course they will. But the model alone is not the substrate of trustworthy science.

 

The intelligence, and the improved outcomes, will only really begin when the two fundamental parts work well together. The room is the working surface: the place scientists stand together, where data is explorable, methods are visible, and interpretation happens. The substrate is what makes that room scientific: governed data, provenance, entity mapping, and a record of what was decided, on what evidence, and by whom.

 

Plenty of tools can give a team a shared space to talk over a chart. What they cannot give you is warrant: the auditable trail that lets a claim made in the room survive contact with someone who was never there. A room without that substrate is just a faster way to agree. A room built on it is a faster way to know why you might be right, where you might be wrong, and what evidence the claim is standing on. That is the room.

 

And I think the organizations that build these rooms well will do more than adopt AI. They will dramatically change the speed of how scientific interpretation happens.

The path ahead

I don't think anyone can claim to know exactly where this goes. Biology has a way of humbling every technology story we tell about it.

 

But I do believe we are underestimating what becomes possible when scientific interpretation moves from static reports into shared, AI-supported environments. I believe we will ask larger questions. I believe we will waste less insight. I believe larger, more diverse teams will be able to participate in complex data analysis. I believe the distance between data generation and biological understanding can shrink dramatically, if we build the right trust architecture underneath it.

 

And I suspect that architecture gets more pressing, not less, as the systems get better. You can see the shape of it already forming across mathematics, where machines are beginning to produce results that can be verified without being followed - checked line by line, and perhaps not understood by the people checking. What's striking is that mathematicians are not satisfied by this. Correctness is not the same as comprehension, and the field keeps working until it has both. Biological science will very likely face the same split. When a system reasons across more evidence than any of us can hold, auditability is what lets us trust the claim - but it is not what lets us understand it. That is the harder job, and it is the one a room is for.

 

This future is the one I feel optimistic about. Not blindly, and not because the models are magic, but because science has always advanced when better tools changed what scientists could see, what they could measure, and who they could work with. AI has the potential to change all three.

 

The first wave has been about individual productivity: writing code faster, summarizing papers faster, automating repetitive work. Every new interface begins with familiar behaviors. The next wave will be about shared scientific interpretation - whether a team can understand complex data faster, whether method choices become more visible, whether claims can carry more of their warrant with them as they travel.

 

That last one is the question I want to keep asking. Can we make confidence travel at the same speed as computation?

 

Not AI as a shortcut around science, but rooms where scientists are not replaced by intelligence and instead surrounded by it. Where human purpose still sets the direction, and every available intelligence can be brought to bear on the problems we most urgently need to solve.

 

The future of AI in science will not be defined by who gives the best prompt. It will be defined by who builds the best rooms.