Cases
The first card is the company: an AI-operated gaming marketplace, still running, with a public write-up rather than an investor deck. Two of the rest are real HTML, CSS and JavaScript lifted out of the repository and running here on anonymised fixtures. The two arena boards were built here from published scientific catalogues, and each one says so on its own dateline.
Most of these are operating systems: something arrives faster than a person can handle it, so a machine researches it, scores it against a written rubric, and hands a human the decision rather than making it. The claim ledger at the bottom is personal research, not a job — the same method pointed at evidence instead of applicants.
We finished the move from a services marketplace to a personal AI assistant for gamers. Three humans and a fleet run it. OrderHub is the OS — buyer chat, catalog, escrow. The marketplace is one of the things the agent can execute. Lifetime scale is the YC page’s own claim, not a current run-rate.
No investor-deck revenue, traffic or unit-economics figures are published here. The checkable numbers live on Y Combinator’s Legionfarm page and in the TechCrunch rounds.
At AGI House I built the first system in the house’s history to research and score the whole community — 13,619 people across six role dimensions, on an n8n agent pipeline with a checker agent behind every researcher. The same database ran the door workflow, the sponsor reports, and an investment view that traces attendees’ companies from an estimated $8.2B at their first event to $55B now.
The scoring numbers are production-scale; the interactive explorer on the case page runs on an anonymised 312-person fixture — the page keeps the two populations apart on purpose, and no per-person record from the private corpus appears anywhere.
For five months, as a contractor, I ran ManyChat’s affiliate program end to end and built the first agentic growth engine in the company’s history — five production AI agents on top of PartnerStack and ManyChat’s own ops systems: 1,900+ applications researched in minutes instead of days, dormant partners re-engaged against their own numbers, the whole active book on a standing quality watch, partner chat up 25×. The channel turned upward the month V1 went live and compounded every month after, ManyChat reached #1 Trending on PartnerStack Discover — and the system stayed in production after the handoff.
The four contract months are on the case page as an index, so you can see the December turn and the compounding yourself — without a single internal dollar published.
A claim about a place on Mars is almost never checked against the place — it is checked against another claim. Mars Arena puts both populations on one rotatable globe with real orbital imagery: the cave and lava-tube candidates NASA and USGS catalogued as somewhere a mission could shelter, and the places American Alchemy named on air as possibly built. One switch moves between them, and the two never blur, because their evidentiary status is not the same. Every marker in the second set carries the established scientific reading in the same card as the claim.
No backend is the point, not a shortcut: the page asks NASA Trek for basemap
tiles and the USGS Astrogeology WMS for each close crop directly from the browser, and both
answer access-control-allow-origin: *, which is what makes a static host
enough. Nothing here is a screenshot of a product — the globe, the table and the imagery are
the working thing.
Eighteen lunar pits from the LROC Pit Atlas, scored on four axes. The board carries a second scoring mode whose only job is to show how much of the ranking rests on measurements nobody has taken: one reading averages the axes that exist, the other charges a pit for the axes that do not. The distance between them gets its own column.
Switching the reading reorders all but the bottom two, and the direction is not random — both fully-measured pits rise, all six two-axis pits fall. The two that rise are the two with published Diviner thermal data, which is also what the literature treats as the type examples. That agreement was not tuned for; it falls out of refusing to score a missing measurement as a zero.
The claim ledger, the disclosure corpus, the graph and the atlas were taken off this site on 30 August 2026. The Mars Arena globe is still here: mars.ab7.ai.
Every one of them is a scoring function someone can read, applied to a working set someone can count, producing an artifact a human then acts on. None of them decide anything by themselves, and that is not modesty about the technology — it is the design. A system that auto-approves is a system whose mistakes you find out about from the person they happened to.
The other shared property is that each case states what is real and what was rebuilt. Lifting an interface out of a repository and running it on invented data is honest; letting a reader believe the invented data is a customer list is not.
Where a case shows a confidence band, the band is derived rather than asserted: eight components are scored separately and summed, so a reader who disagrees can disagree with one line instead of with a verdict. Each case documents its own rubric on its own page.
Client and partner work with no publishable artifact is not listed here or on the overview. A credential I cannot hand you a document for is not a credential.