The verification layer for AI for Science

SwarmLabs replaces wet-lab experiments with virtual ones and issues auditable 10-chapter V&V reports (aligned to ASME V&V 10-2019 and the FDA CM&S outline). The scarce asset is not another model that generates hypotheses — it is honest uncertainty quantification that tells you when a prediction can be trusted, and when it cannot.

62
validated scenarios
(published, none filtered)
157,768
structured scientific
entities (works/authors/concepts)
62 / 0 / 0
V&V verdicts
PASS / MARGINAL / REFUTED
10
chapters per report
ASME V&V 10-2019 aligned
MIT
open-source SDK
swarmlabs-engine-kit

Why this layer, and why now

The measurement / evaluation layer of AI for Science has become strategic infrastructure — and it is exactly the layer that model-first labs are missing.

💰 Capital has repriced verification

UniPat AI raised $300M at a $2.5B valuation (Alibaba led, Tencent participated) to build synthetic data and AI evaluation / benchmark design. Anthropic acquired Coefficient Bio for $400M to bring rigorous verification into its science stack. Whoever measures others' models holds the leverage.

🧪 Generation is solved-ish; trust is not

Foundation models can now propose hypotheses. Almost none can tell you, with a computed number, how confident a result is or where the model stops being valid. SwarmLabs is that missing layer: GP surrogates + coverage-audited intervals + an explicit out-of-distribution guard.

📄 Deliverable, not a chat log

Every scenario produces an archived 10-chapter V&V report with R², single-sided coverage, variance-calibration κ, and a reproducibility manifest (seeds / bounds / re-run command) — exportable to HTML, JSON and DOCX. Chat products cannot archive a defensible audit trail.

🔁 Prompt-to-validation, and a self-evolving loop

Where the field says "Prompt-to-Drug", the gate is the prompt-to-validation step that hunch-driven pipelines skip. Ours runs daily: a signal→candidate→verify→promote→effect flywheel where every promotion clears the gate above — recursive self-improvement with an audit trail, not a claim.

The GPT-6 moment — even frontier models need a gate

In 2026 a general-purpose research model topped a drug-developability benchmark with no bioinformatics plugin and no fine-tuning. What the reviews said next is the whole thesis of this company.

📈 What happened

GPT-6 Astra placed first on Insilico Medicine's DDD antibody-developability benchmark (37.98) — no plugins, no fine-tuning — while general biomedical agents (Biomni, DrugAgent) closed in on expert baselines. Generation is no longer the scarce input.

⚠️ What the reviews say

Practitioner reviews of these agents converge on one failure mode: they are "prone to plausible-but-wrong outputs, weak guardrails." The prescription is explicit — a held-out benchmark + quantified uncertainty + a human checkpoint.

🛡️ Where SwarmLabs sits

That prescription is SwarmLabs. We do not compete to out-predict the frontier model; we are the layer that says whether a given prediction can be trusted. The stronger the base model, the more that layer is worth.

And it is callable, not a slide. The engine is exposed as an MCP server (verify_predictionPROCEED / BLOCK_AUTONOMOUS_ACTION) and as a read-only HTTP gate (GET /v3/gate/{key}) any agent can query without running Python. Confidence below the calibrated threshold returns an explicit block — never a confident guess.

What we do differently

Honest positioning: we are not "another AI science assistant". We are the verification substrate beneath them.

DimensionChat-based AI science toolsSwarmLabs
Core valueRead / write / summarize the literatureReplace real experiments (GP surrogate + UQ + virtual experiments)
RigorCitations & summaries; no quantitative validity proof10-chapter V&V report; coverage / κ calibrated against independent sets
UncertaintyNo numeric UQ (never answers "how trustworthy is this?")3% noise floor never lowered; single-sided coverage verdicts; κ only loosens
ReproducibilityConversation output; hard to archive62 reproducibility manifests (seeds / bounds / re-run command)
FailuresSystematically filtered / undisclosedDisclosed — held-out REFUTED cases kept and explained, not hidden

Assets you can inspect right now

Everything below is live and machine-readable — please verify rather than take our word for it.

Where we are (unsanitized)

⚠️ Honest status — please read before any conversation

  • Pre-revenue: no paying customers yet. The engine is live; the license tier (Pro / Lifetime) is early.
  • Single maintainer (bus factor = 1). This is a real institutional-funding obstacle, stated up front.
  • No corporate entity yet — it gates most accelerator / cloud-startup channels.
  • Reports are validated against analytical solutions and published benchmarks, not yet against our own wet-lab runs. Parity against real measurements is the next milestone.
  • 0 of the 62 validated scenarios are REFUTED in the main ledger. And we do not manufacture a clean scoreboard: 8 are REFUTED under robustness held-out stress and 53 under extrapolation stress, and those failures stay visible on the ledger.

Talk to us

Warm intros, accelerator programs, and strategic collaborations welcome. If you are building AI4S and need a verification substrate, we would like to hear from you.

[email protected]