REGISTERED CALL C4
Will a model that runs offline on a laptop or phone match the top Artificial Analysis score of Oct 2, 2026 by Apr 2, 2027?
Claim. On or before 2027-04-02, a model that runs fully offline on one consumer laptop or phone (up to 64 GB of memory, no network use at inference) scores at or above 58 on the Artificial Analysis Intelligence Index v4.3, the score of the #1 model on that index on 2026-10-02 (Claude Opus 5.5, max with fallback).
Resolution rule. Resolves YES if, on or before 2027-04-02 23:59 Pacific Time, Artificial Analysis (artificialanalysis.ai) publishes an Intelligence Index score of 58 or higher (on index v4.3, or on a later version where the 2026-10-02 leader, Claude Opus 5.5 max with fallback, is re-scored: then the leader's re-scored value) for a model whose weights are publicly downloadable and which has a documented run fully offline on one consumer laptop or phone with up to 64 GB of memory. The documented run may be by Artificial Analysis or a public, reproducible run by anyone. Resolves NO otherwise. Not counted: benchmark results published only by the model's maker, models offered only through an API, quantized variants that Artificial Analysis has not scored, and any index other than Artificial Analysis. The 2026-10-02 reference page is archived at https://web.archive.org/web/20261003030249/https://artificialanalysis.ai/leaderboards/models.
Clarification, 2026-10-03
Clarification of C4. The registered text is unchanged and governs; this entry says how it will be applied.
1. Threshold. Artificial Analysis publishes Intelligence Index scores as whole numbers. The 2026-10-02 leader's published score was 58 (archive link in the registration). The threshold is the number Artificial Analysis publishes: 58. If Artificial Analysis later publishes unrounded scores, the candidate must reach at least 58.0 unrounded on index v4.3. A mirror site lists 57.6 for a Claude Opus 5.5 configuration; that is not an Artificial Analysis source and is not used.
2. Offline. The full model weights are on the device before the run starts. The whole inference runs on that one device (laptop or phone, up to 64 GB). No cloud fallback, no fetching of remote weights during the run, no hosted tools, no remote retrieval, and no routing to another machine at any step.
3. Same protocol and effort. Artificial Analysis scores every model with the same ten evaluations; each model is scored in a named configuration with a reasoning effort setting. The leader was scored at max effort (Adaptive Reasoning, Max Effort, Default Fallback). The candidate counts at the configuration Artificial Analysis scored, and the offline run must use the same weights and the same configuration, including the same reasoning effort. A score at one effort does not count for a run at another.
4. Who scores. Artificial Analysis benchmarks hosted endpoints and chooses which models to score; SOTO has no way to request a score. A candidate Artificial Analysis never scores cannot make C4 resolve YES.
Clarification, 2026-10-03
Second clarification of C4. The registered text is unchanged and governs. This entry settles how the registered rule and the clarification in chain record 7 fit together.
1. Index version. While Artificial Analysis publishes Intelligence Index v4.3 (including its patch versions, v4.3.x), the bar is the fixed value from the first clarification: a published score of 58 or higher, or, if unrounded scores are published, at least 58.0. The registered rule's re-scored path applies only if Artificial Analysis retires v4.3. In that case the bar is exactly the 2026-10-02 leader's (Claude Opus 5.5, max with fallback) re-scored value on the successor version, as Artificial Analysis publishes it.
2. Fallback models. If the candidate uses any fallback model, that fallback model and its weights must also be fully on the device before the run starts, and all inference, fallback included, runs on that one device. No remote fallback at any step.
Integrity checks
- Resolution sources named: https://artificialanalysis.ai/leaderboards/models; https://web.archive.org/web/20261003030249/https://artificialanalysis.ai/leaderboards/models; https://conway.tech/.
- Exclusions stated: Benchmark results published only by the model's maker.; Models offered only through an API.; Quantized variants that Artificial Analysis has not scored.; Any index other than the Artificial Analysis Intelligence Index..
- Ambiguity test at registration: 20, 20, 20 of 20 readings agreed per scenario (chain record).
- Resolver (19 of 20 readings must agree): runs after the deadline, not run yet.
Estimate. SOTO’s estimate is 8%; the obvious estimate is 35%.
Judgment on the obvious: this call never counts toward SOTO’s beats-the-obvious score (rule 10).
Deadline. 2027-04-02T23:59:00-07:00
Exclusions
- Benchmark results published only by the model's maker.
- Models offered only through an API.
- Quantized variants that Artificial Analysis has not scored.
- Any index other than the Artificial Analysis Intelligence Index.
Sources that settle it
- https://artificialanalysis.ai/leaderboards/models
- https://web.archive.org/web/20261003030249/https://artificialanalysis.ai/leaderboards/models
- https://conway.tech/