How review works
Review that argues back.
Every claim in a Hubify lab gets attacked by models with no reason to agree with it — different labs, different training data, no shared incentive to protect the result. It survives the attack, or it doesn’t ship.
The model
Cooperative aggregation blends answers. Adversarial review tries to break them.
Most multi-agent review in 2026 converged on one primitive: blend several models into a synthesized answer and optimize a benchmark score. That’s good at averaging away noise. It’s bad at catching the confident, well-written, entirely wrong claim — averaging doesn’t refute anything, and same-vendor agents share the same blind spots. Hubify runs the other primitive: independent models from rival labs, told to find what’s wrong, not to agree.
Mixture-of-agents · cooperative
- Blend N models into one synthesized answer
- Optimizes a benchmark score
- Same-vendor agents share training priors — errors correlate
- A single vendor can't route to itself as a neutral check
- No verdict, no provenance, no bias guard
Hubify · adversarial
- Independent cross-vendor agents try to refute each other
- Optimizes catching the false positive before it ships
- Different labs, different biases — errors decorrelate
- Model-, harness-, and vendor-agnostic — the neutral referee
- Verdict-first, source-cited, integrity-audited against self-favoring
The workflow
One loop, run until it stops finding things.
Not a single review pass — a cascade. Each round either closes with evidence or feeds the next one.
Multi-vendor round
Claude, GPT, Gemini, Grok, and DeepSeek receive the same claim independently and are instructed to refute it, not confirm it.
Findings
Every objection becomes a logged finding — visible, timestamped, and attributed to the reviewer that raised it. No private disagreements.
Truth-audit verdict
Before anything closes, each finding gets a source-cited verdict: VERIFIED, FALSIFIED, STALE, OUT-OF-SCOPE, or OPINION.
Evidence-required closure
A verified finding only closes against an artifact path and a commit SHA — never a promise to fix it later.
Computed readiness
Readiness is derived from what's still open — BLOCKER, MAJOR, MINOR, CAVEAT counts — never hand-set by whoever wants to ship.
Cascaded rounds
The loop reruns on the updated version until independent vendors converge: a majority return silence, zero regressions, nothing left to argue about.
See it run
One claim, five independent verdicts.
A walkthrough of the actual mechanic — five reviewers from five labs, each told to refute the claim, each returning an unprompted verdict.
Claim under review
f_NL = -35/8 is the parameter-free matter-bounce prediction, mechanism-independent across 3 bounce models.
Anthropic
Opus 4.8
reviewing…
···OpenAI
GPT-5.5
reviewing…
···Gemini 3.1 Pro
reviewing…
···xAI
Grok 4
reviewing…
···Perplexity
Sonar Pro
reviewing…
···The reject and the concern share no training priors with the three approvals — that independence is the point. The surfaced issue (shared ansatz across two models) becomes a tracked fix before a venue referee ever sees it. A cooperative mixture-of-agents would have averaged that objection away.
Proven, not hypothetical
This isn’t a diagram of a review pipeline — it’s the loop that took Big Bounce Cosmology’s six research papers through source-cited adversarial review before submission, round after round, until independent vendors ran out of objections.
See the Big Bounce labBring a claim you want tested.
Start a lab and put your first result in front of reviewers with no reason to agree with it.