Every AI answer, cross-examined by rival models before anyone acts on it. You get the record: what survived, what was killed, and where they still disagree.
or try the live room yourself
Chapters, papers and position documents go through the room before reviewers and readers see them. The findings come back ranked by a printed support score: how many models found the same thing independently, and whether it survived challenge.
Draft legislation, legal-basis analyses, consultation responses. Rival models attack the argument through separate lenses, and the dissents are kept instead of being averaged away.
One API call between a model’s output and the people or systems that act on it. Rival vendors’ models cross-examine the answer, and the full record of who said what comes back with it.
A printed support score built from what the room did: independent blind discoveries, survival under challenge, recurrence. Arithmetic, never model confidence.
Findings the room itself took down in cross-examination. The final text is checked mechanically against the killed list.
Where rival models still disagree at the close, the report states both positions and says so: this is your call, not the room’s.
The checklist: wrong cross-references, names that drift, numbers that do not match.
The lineup as run, rounds, lenses, and the caveats stated against ourselves, including what the instrument cannot tell you: whether a finding is correct.
Your source document is purged after 30 days. Press erase on the report page and everything goes at once.
Every seat answers alone. No model sees another model’s answer, so early agreement is never imitation.
The seats read each other’s answers and challenge or endorse them. A seat may hold, move or concede, and every move is recorded.
Findings are ranked by the support score: how many seats found the same thing blind, whether it survived challenge, where it recurred. Remaining disagreement is reported as disagreement.
When models from rival vendors agree, is that evidence, or a blind spot they share? We measure what is left of an agreement after structured challenge.
Motivated closure, plan-continuation bias and context drift: three failures that never look like a mistake from inside one conversation. Followed by a live four-model demo.
5,000 GPU-hours on the Discoverer supercomputer in Sofia, August to November 2026, for running open-weight models as referee seats next to the commercial frontier models.
The consumer product is where the method runs in the open. One prompt, four AIs, one thread, no edits.
25 years in software: seven years on mission-critical lithography software at ASML, user-system research at Philips, then his own studios. MSc Computer Science, PDEng TU Eindhoven.
LinkedInSix years in finance at wagamama US in Boston, from general accountant to finance manager. Northeastern University. Owns Collider’s cost model and unit economics.
LinkedInCollider started at home. When our son was born we kept asking ChatGPT, Claude and Gemini the same questions, and each gave a confident, different answer. So we put them in one room and made them answer to each other. Then institutions started sending us documents, and the referee pipeline grew out of that.