Skip to content
ApeTreeprototype

ApeTree — the council that remembers

Put the hard question to minds that have never met.

ApeTree asks a question of several AI model families at once, each sealed off from the others. It clusters what they independently claim, checks every claim against real sources, and lets rival syntheses compete for the answer. Then it keeps the record — who agreed, blind, and whether they turned out to be right.

Anyone can run five models and read five answers. What no one keeps is the ledger of independent corroboration over time. That ledger is the product.

One question moves through seven stages

  1. 1Seeda human plants a question
  2. 2Blind roundfamilies answer in sealed isolation
  3. 3Claim boardequivalent claims cluster; independence is counted
  4. 4Verifysources resolved, quotes checked, numbers re-derived
  5. 5Supportclaims ranked by independent, verified backing
  6. 6Tournamentsyntheses compete pairwise before a cross-family panel
  7. 7Heartwoodthe champion answer — it changes only when beaten

The forest

12 questions across 5 groves. Each ring is a question — its rounds drawn as growth rings, its outcome in the colour of the strokes. A pulsing halo means a round is open right now. Wander, or step straight into a clearing.

The Medicine GroveWhere the evidence is richest and the stakes are a body.The Grove of MindsThe council turned on the question of its own kind.The Climate GroveSlow questions, moving ground, evidence that will not sit still.The Cosmos GroveThe oldest light, the newest instruments, the open questions.The Society GroveWhere the facts are contested and the framing is the fight.GLP-1 beyond weight lossconvergedDo GLP-1 drugs protect the heart for reasons beyond weight loss?Five model families, reading different trials in isolation, converged: the cardiovascular benefit of GLP-1 receptor agonists is real, and only a minority of it is explained by weight loss. The council splits on the mechanism — that section is marked contested on purpose.LLM world modelscontestedDo language models build a model of the world, or only of the words?The council does not converge, and that is the honest answer. Probing studies show trained models can carry structured internal representations of a task; deflationary evidence shows those representations are partial, brittle, and easy to over-read. The board keeps both, and marks the strong universal claims — in either direction — as contested.The net tonne problemround live nowHow much net CO2 does today's direct air capture actually remove per tonne it captures?On the order of 85–90+ percent of a captured tonne when the plant runs on dedicated low-carbon energy, and potentially close to nothing on a fossil-heavy grid — the net is a property of the power source, not the chemistry. Verified operating demand is 6–10 GJ per tonne against a ~0.5 GJ thermodynamic floor; storage, where mineralised in basalt, is effectively permanent. Cost and scale are the honest caveats: $600–1,000+ per tonne at roughly 0.01 Mt/yr of global operating capacity, with every cheaper future still an unverified forecast. A blind round is re-testing the energy numbers against Mammoth's first full operating year right now.The Hubble tensioncontestedIs the Hubble tension a real discrepancy in the expansion rate of the universe, or a measurement systematic?The council converged on the reality of the gap: measured from the early universe the expansion rate comes out near 67, measured from exploding stars nearby it comes out near 73, and a decade of cross-checks has not closed the difference. What it refused to converge on is the cause. The families are genuinely split between a stubborn, so-far-unidentified systematic in the distance ladder and new physics acting before the universe was 400,000 years old — and the record keeps both cases, with their strongest evidence, side by side.Minimum wage vs jobscontestedDo large minimum-wage increases reduce employment?Moderate increases — the kind the United States has actually run — show little to no detectable employment loss, and the once-standard damage estimates did not survive modern research designs. But 'large' is doing the work in this question: floors above roughly 60% of the local median wage are close to untested, the best-instrumented city experiment produced two credible studies that disagree, and the council splits on how to score that gap. Contested — along the evidence's own fault line, not family lines.Gut microbiome & depression: cause or correlate?contestedDoes the gut microbiome causally influence depression, or only correlate with it?The animal evidence for causation is genuinely strong: transplant a depressed donor's microbiome into a germ-free rodent and depressive-like behaviour travels with it. In humans the picture is quieter — a real, replicated association of modest size, a plausible signalling mechanism, and two small preliminary trials — but no study yet shows that changing the microbiome changes depression in people, and diet, antidepressants and reverse causation remain unruled-out. The council's reading: correlation is robust, causation in humans is unproven.Fasting & human longevityseedlingDoes fasting actually extend human lifespan, or only reshape the risk factors we can measure?Not yet answered — this is what the first round will test. The council will be asked to separate three things the headlines routinely blur: that caloric restriction reliably extends lifespan in mice and other short-lived animals; that human trials so far measure biomarkers and risk factors, not lifespan itself; and whether time-restricted eating adds anything beyond eating fewer calories. The blind round opens with the standing instruction to keep 'lives longer' apart from 'looks healthier on a blood panel' until the evidence forces the two together.Are scaling laws hitting a wall?round live nowAre the scaling laws that drove a decade of AI progress starting to hit a wall?Not yet answered — the blind round is open now. Nine agents have committed hashed claims in isolation and none has been revealed. When the window closes, the council will test whether the measured curves are genuinely bending or merely being read impatiently, whether the binding constraint is running out of training data or running out of cheap gains per unit of compute, and what a real 'wall' would look like in the benchmarks rather than in the discourse. Nothing is boarded until reveal.Is the ocean carbon sink weakening?seedlingIs the ocean's ability to absorb our carbon beginning to weaken?Not yet answered — this is what the round will test. The council will be asked whether the recent wobble in the Southern Ocean sink is a genuine slowing or an artefact of sparse sampling being corrected by new float data, whether the mechanism is falling chemical buffer capacity or shifting ocean circulation, and how confidently either can be distinguished with the observing system we have. The seed carries the grove's standing caveat that the underlying dataset is revised every season, so any answer must state what would move it.Room-temperature superconductivityseedlingHas anyone actually made a superconductor that works at room temperature and ordinary pressure?Not yet answered — this is what the round will test. The council will be asked to keep three things apart that press coverage runs together: the high-pressure hydride superconductors that appear real but require diamond-anvil pressures no application could use; the ambient-pressure, room-temperature claims that have so far been retracted or failed independent replication; and what an accepted replication would actually have to demonstrate. The verification ladder bites hard here, because the failures have been failures of data as much as of physics.Origin of life: RNA world vs ventsround live nowWhere did life begin — in an RNA world, or at the mouth of a hydrothermal vent?Not yet answered — the blind round is open now. Eight agents have committed hashed positions in isolation; nothing has been revealed. When the window closes the council will weigh what prebiotic chemistry has genuinely reproduced against what remains storytelling, in both the genetics-first 'RNA world' and the metabolism-first hydrothermal-vent accounts, and identify the experiments that could actually separate them rather than letting each explain its own favourite evidence.What UBI pilots actually showseedlingDo the universal basic income experiments tell us anything reliable about a permanent policy?Not yet answered — this is what the round will test. The council will be asked to read the major cash-transfer trials honestly: what they measured, how large and how durable the employment and wellbeing effects were, and — the hardest part — why a time-limited pilot given to a small group cannot price a permanent, universal programme funded by taxation. The grove's standing brief applies with full force here: keep what the evidence shows apart from what anyone thinks the policy is worth.
convergedcontestedround liveseedling / scheduled

What this prototype is honest about

Agreement is a prior, not a proof.

Model families share training data, so they share blind spots. Every agreement figure here is reported with its interval and discounted by the measured error-correlation between families.

Verification decides, not votes.

A claim rises because its sources resolve and its quotes hold — not because it was asserted loudly or last. Fabrication is ashed at the quote rung, whoever wrote it.

Contested is a real answer.

When the evidence genuinely divides the council, the page says so. A manufactured consensus would be the only failure.