Login

Benchmarks

Ten diligence modules, one target, four systems

VentriLinks Diligence against ChatGPT Deep Research, Claude + Research, and Gemini Deep Research across ten diligence modules on a live target.

Read the method

Standings

#SystemTotal /500Avg /50Module winsValid findingsWrong findingsClaims enumeratedVerifiedFalse claims
1
VentriLinks Diligence+105.3
478.847.91016352,50498.6%29 (1.16%)
2
ChatGPT Deep Research−105.3
373.537.40124161,60392.8%93 (5.84%)
3
Claude + Research−108.4
370.437.00148191,85294.2%87 (4.73%)
4
Gemini Deep Research−194.1
284.728.5066271,11085.7%100 (9.35%)

Module by module

Module 01

Scientific validation

Construct, mechanism, and evidence grade for the lead asset, traced through primary literature with effect sizes and reproducibility checks.

Against the field · /50

01020304050
  • 48.4+0.3
  • 48.1−0.3
  • 43.6−4.8
  • 27.8−20.6
  • Diligence
  • ChatGPT
  • Claude
  • Gemini

VentriLinks Diligence · axis rates

Fact accuracy
99.7%
False claim rate
0.3%
Source quality
86.0%
Report sections filled
100.0%
Provenance honesty
100.0%
Evidence artifacts discovered
20

Breakdown · normalized scores · all four systems

Rubric axis
Diligence
Claude
ChatGPT
Gemini
Fact accuracy/1010.09.910.09.2
False claim rate/109.89.59.86.2
Source quality/108.64.89.15.5
Report sections filled/1010.010.010.02.9
Provenance honesty/1010.09.49.24.0
Evidence artifacts discovered20412418

Method

Which target does the benchmark use, and why this one?

Every run answers about one asset: FT819, the off-the-shelf, iPSC-derived CD19 CAR T-cell candidate of Fate Therapeutics (NASDAQ: FATE). Four systems each ran the ten modules on it, which gives the forty runs on this page.

The target is public, so SEC filings settle the financial claims. The asset carries a real patent lineage — the 1XX CAR construct and the TRAC-locus insertion out of Memorial Sloan Kettering — so the freedom-to-operate module has ground truth in the patent registers.

Two facts about the program changed before the runs. The lead indication moved from B-cell malignancy into autoimmune disease, and the Janssen collaboration ended. Each change measures whether a system reports the old record as current.

Which prompt did each system receive?

One prompt per module, and every system receives the same text. Each prompt derives from the module skill the product itself runs: module 01 comes from SCIENTIFIC_VALIDATION.md, module 02 from FTO_ANALYSIS.md, and so on for all ten.

The derived prompt drops what an outside system cannot have. The internal block schema becomes plain markdown tables, the knowledge-graph and tool names go, the cross-module handshakes go, and the references to an uploaded data room go. The prompt never names VentriLinks.

What remains is the analysis task: the numbered workflow, the risk roster with its five columns, the mitigation table, and the sourcing rule. That rule asks for a primary source with a retrieval date on every factual claim, as an inline link, and for a statement of which sources the system searched and where its coverage is incomplete.

Each run starts in a fresh chat with deep research and every connector on. There is no follow-up turn and no repair turn.

What is the scoring procedure for one module?

Five axes: fact accuracy, false claim rate, source quality, report sections filled, provenance honesty. The grader records a numerator, a denominator and a rate for each one, and the score follows by arithmetic. A module is /50 and a run is /500, to one decimal place.

Four axes score 10 × rate. The false-claim axis maps differently, because a perfect run scores zero on it: it maps 0% error to 10 and 20% error to 0, and stays at 0 below that.

A gate failure caps its axis at 4.0, whatever the counts say. An invented patent number, a legal status taken from an aggregator rather than a register, a runway with no arithmetic, and a discontinued program listed as active are gate failures.

What does the grader check inside each module?

The rubric fixes a deliverable list per module, and report sections filled scores populated deliverables over that list. Module 02 has seven: blocking patents in risk tiers, chain of title, inbound licenses, jurisdictional summary, legal-status verification, validity exposure, and the risk roster. An empty table, a table with the wrong columns, or a paragraph that dodges the question does not count.

Each module also names the claims the audit reaches first, where the modules differ from each other. On module 02 that is every patent and application number, the assignees, the expiry dates, and the legal status. On module 04 it is the registry facts — phase, arms, enrollment, endpoints, dates — against the registry record. On module 09 it is the runway arithmetic against the filings.

Provenance honesty scores a separate unit: each source-class label, each declared gap, each unverified mark, each derived-or-assumed flag, each data cutoff, and each search-scope statement. It fails when the label is missing, the label is wrong, or a source contradicts a stated absence.

What counts as one claim, and how was the census enumerated?

One claim is one assertion about the world that a primary source can settle. The test is simple: if a primary source could prove it wrong, it is a claim; if disagreement with it is a matter of opinion, it is not. A recommendation, a forward-looking judgment, a severity rating, and a verdict on evidence the run already stated are not claims.

A table row usually carries several. “RECLAIM-LN is Phase 2, open-label, single-arm, n=53” is four claims.

The enumerator reads the run in document order and writes one numbered line per claim, with the verdict, the basis, and the propagation. A fact restated in five deliverables is one claim, and the propagation column records the restatement. There is no cap and no quota — 7,069 claims across the forty runs.

Each line takes one of five verdicts: true, false, unsupported, stale, unsettled. Fact accuracy is true over settled claims, and the false-claim rate is false over settled, where settled is true plus false plus stale. Verification coverage — settled over enumerated — is reported beside every module and never scored.

How does the audit define a verified claim and a false claim?

Verified is the share of enumerated claims a source confirmed correct. It is reported and never scored.

False claims counts only plainly wrong statements. A correct fact with no source is a sourcing failure, charged once to source quality, not a false claim. Source quality is therefore scored over citation instances plus unsourced claims, so a run cannot score well on it by citing less.

How does the audit define a valid finding, and why is there no recall figure?

A valid finding is a real, correctly-stated item: a patent family with the right status, an NCT that matches the registry, a financing event that matches the filing. Each item counts once — one patent family, not each member; one competitor program, not each trial.

We do not compute recall. Its denominator cannot be enumerated for these modules, and a capped denominator measures the grading budget, not the model.

What do the planted traps measure?

Six traps sit in the target’s public record. Diligence passes five and fails the legal-status trap: it printed “Granted” for a European patent the EPO had revoked.

The lead-indication trap separates the field. Gemini reports the discontinued oncology indication as live in five of its ten modules and builds a $1.78B market on it; Claude and ChatGPT pass it in all ten.