INSIGHTS
Article

Governing AI Use in Life Sciences R&D

In life sciences R&D, scientists already use AI, whether or not anyone formally approved it: built into workflows, running against lab data, often without a sign-off process to speak of. We’re seeing it everywhere we work in the sector, sanctioned or not.

What’s still being worked out is what happens next: the space between a scientist’s raw, AI-assisted result and something the organization can actually trust, reuse, and trace back to its source. That gap, not the technology itself, is where the risk lives. It’s also where the real differentiation between organizations is starting to show.

The problem: productivity without a paper trail

The failure mode looks similar wherever it shows up. A scientist spins up an application, feeds it raw data, and gets a useful result, but there’s no record of how the tool produced it, no way to validate the output, and no owner once the person who built it moves on. Multiply that pattern across a research organization, and what you get isn’t AI-enabled productivity so much as technical debt accumulating quietly: an expanding surface of undocumented, unowned software and analysis. Token costs climb alongside it, and teams are often left holding codebases they can’t read or maintain.

A scientist may use AI to summarize assay results, generate code for data cleanup, compare compound profiles, or explore hypotheses across experimental datasets. Each use is helpful in isolation. The problem is that the logic, assumptions, prompts, data extracts, and validation steps often disappear along with the individual workflow that produced them.

This isn’t hypothetical. Cloud Security Alliance’s 2026 research on shadow AI and orphaned agents cites Strata’s 2026 AI Agent Identity Crisis findings: only 23% of organizations have a formal, enterprise-wide strategy for agent identity management, a sign of how immature ownership and control structures remain.

Locking AI down is the understandable response, but it tends to backfire: restrict access, and teams lose the productivity gains that justified the tools, often routing around IT. Leave AI unrestricted, and the result is the fragmentation above, only worse: not just data scattered across spreadsheets, the classic informatics headache, but software, models, and analytical logic proliferating with no governance at all.

The organizations getting this right aren’t choosing between those two options. They’re building a third path: centralizing and enabling AI through a governed platform, rather than restricting it or leaving it alone. That means keeping data centralized, giving scientists a set of approved tools to build with, and creating a controlled route from a working prototype to something the organization can deploy, reuse, and audit with confidence.

Three practices stand out as genuinely effective at closing that gap. Each applies at a different point in an AI system’s life: deciding whether it’s worth building in the first place, checking it before it reaches production, and reviewing it once it’s live.

Before building: choose the right level of AI

The first move happens before any building starts: deciding whether a use case needs AI at all. It’s tempting to treat AI deployment as free once the tools exist, but it isn’t: every agentic build carries validation and maintenance costs, a simpler approach might not. The most disciplined research groups ask, case by case, whether a problem genuinely requires a full agentic system, or whether a well-scoped one-shot LLM call or a simpler retrieval pattern would do the job at a fraction of the cost and risk. That trade-off is only getting sharper as token economics becomes a real, compounding cost rather than a rounding error, a live concern for AI leaders across the industry right now. Managing it starts upstream: by not overbuilding in the first place.

A 2026 piece in Clinical Pharmacology & Therapeutics, co-authored by a scientist inside a major pharmaceutical company’s drug metabolism and pharmacokinetics group, makes the same case at an industry level: organizations need a per-application cost-benefit framework for AI delegation, rather than defaulting to it whenever it’s technically feasible. The efficiency gains from handing a task to AI, the authors argue, may not offset the risk and validation costs it introduces, and making that trade-off explicit, case by case, is becoming part of how mature research teams operate.

Before deployment: validate before handoff

The second move happens once a prototype exists but before it reaches production: a formal checkpoint where the work gets pressure-tested ahead of handoff. Some of the most effective research-informatics groups have built exactly this into their operating model: a defined step between a scientist’s working prototype and its release to the broader organization, where it’s tested against real conditions before anyone downstream relies on it.

The role isn’t to gatekeep discovery itself. These teams typically pick up after the science is done (after a scientist or discovery group has already produced a model or analysis), and their job is to get that output into usable, reliable form faster, not slower. Done well, this kind of governance speeds up trustworthy deployment; it doesn’t just add friction.

That instinct has regulatory backing too. In January 2025, the FDA issued draft guidance proposing a risk-based credibility assessment framework for AI models used to produce information or data supporting regulatory decision-making for drug and biological products. The framework lays out a seven-step process: defining a model’s context of use, assessing its risk, and documenting evidence of its reliability before its outputs are used to support a regulatory decision. It’s a formal version of the same principle: don’t let an AI-assisted result reach production, or a submission, without a structured check first.

Once live: reviewing outputs on an ongoing basis

The third move isn’t a one-time gate. It’s continuous. Some computational biology teams are experimenting with what they call an “agent council”: a standing structure where AI agents check and challenge each other’s assumptions, reviewed on a regular cadence specifically to catch drift and hallucination before it compounds. It’s a deliberate trade-off between speed and caution, one that’s rarely fully resolved, but the structure means errors get caught by design, not by accident.

The underlying idea, structured cross-checks between AI systems before human review, is supported by research on multi-agent debate. In a 2023 arXiv paper later published at ICML/PMLR, “Improving Factuality and Reasoning in Language Models through Multiagent Debate”, found that having multiple model instances debate and challenge each other’s answers measurably reduced factual errors and hallucinations compared to single-model responses. It isn’t a substitute for human accountability, but it can be a useful design pattern for surfacing inconsistencies earlier.

What comes after adoption

The question in life sciences R&D was never whether scientists would use AI. What’s now emerging is the discipline around it. The organizations getting this right aren’t relying on a single control; they’re building governance in layers across the AI lifecycle, with a thoughtful decision before building, a checkpoint before deployment, and ongoing review once something’s live.

Applied together, these practices compound. Each one closes a different gap, and the combination moves an organization toward a governance state mature enough to capture what AI actually offers: productivity gains without sacrificing traceability, ownership, or trust. That’s the frontier worth watching in life sciences right now: not who adopted AI fastest, but who built the governance to make that adoption durable.

By the Arrayo Insights Desk