THE RECORD
Every call SOTO has made, kept as it was made
Signals are registered before the outcome is known. Nothing is edited after registration, and misses stay on the record.
Registered calls: 4. Verdicts recorded: 0. The record opened Oct 1, 2026, at 00:00 PT.
The chain
SHA-256(prev_hash || canonical_json). Since Oct 1, 2026, the day of the first registered call, software anchors the latest hash outside SOTO every night at 23:59 PT (ADR-000, rule 11). Each anchor is a record in the chain.
- Record 1 GENESIS: hash 10baca9148e0d037b9b905e931933395f6845d9cded5e8bf8d363437bec5ede9
- Record 2 C1: hash 60eb50cfc1032389791de96d5a83544441b756317516563cf5fea78ce8390850
- Record 3 ANCHOR-2026-10-01: hash b240b553c1cd4e719cf98f54758d58c2222fb6a6f91620dff4ebee5f9472006e
- Record 4 C2: hash 8541d9d43f15255498200898c16f6d09dcda8c83e90013e95155ccd589aa5c88
- Record 5 C3: hash 45a3e5096eb1a4171dfc830827186018d2afce0f3558dc4116f3aebed00f27fe
- Record 6 C4: hash a5f03e6dce17dc790f3570cd1e9c7de1fae2cb806c90c14aae3a36877015fc0d
- Record 7 C4-CLARIFICATION-7: hash 635ed38c2d6606ee2f8222066ca8cc9ec47d15729448b628c3b8b02a078fd53e
- Record 8 C4-CLARIFICATION-8: hash 43c26c34cbd4ddd2513880af9073742e01f290f01855f6167c8b54b31e0bc6b7
- Record 9 ADR-001: hash 4e30dc43ee193ea5adab5cb75f5d2d621eabdd3c155d257f547219400e32e8e0
chain.jsonl · readings.jsonl · rehearsal chain
ADR-000: The rules of the record
SHA-256 1bd75421d340cc3e4babbbb448d91283d478b3eec96b08eef5f70f2c98bdb8c7
# ADR-000: The rules of the record Status: accepted. Effective October 1, 2026, at 00:00 Pacific time (2026-10-01T07:00:00Z). These rules apply to every call registered from that moment on. 1. **Calls.** A call is a claim about a future, publicly observable event that SOTO registers under these rules. Only registered calls are scored. 2. **Contents.** At registration, a call records: the claim; the resolution rule, which states the exact condition for YES, with anything else resolving NO; a deadline with date, time and time zone; the sources that will settle it; the exclusions, which say what does not count; the registered probability p; the base rate b; the base-rate type, which is either REFERENCE CLASS or JUDGMENT; for a REFERENCE CLASS base rate, the derivation, which gives the reference class, the sample with links, the count and the time window; for a JUDGMENT base rate, a one-sentence reason and an explicit statement that no reference class is claimed; and the registration time in UTC, to the second. p and b are whole numbers from 1 to 99 percent, and p is never equal to b. 3. **Admission.** Software registers a call only if every check passes: the claim and the rule use none of the words comparable, significant, major, meaningful, notable, large or several; the deadline falls between 1 and 365 days after registration; the sources and exclusions are named; the base-rate type is stated; a REFERENCE CLASS base rate is published with its derivation, and a JUDGMENT base rate is published with its reason and the label JUDGMENT; and in an ambiguity test, 20 independent model readings of the rule, applied to at least three hypothetical outcomes written for the test, reach the same verdict in at least 19 of 20 readings for every outcome. If any check fails, the call is not registered. 4. **Chain.** The genesis record and every registration, update, verdict, withdrawal, correction and anchor are records in one hash chain. Each record is serialized as canonical JSON, with keys sorted, separators "," and ":", UTF-8 encoding and non-ASCII characters left unescaped. Records are numbered from 1, in order of appending. h(n) is the SHA-256, written in lowercase hexadecimal, of the UTF-8 bytes of the 64-character string h(n-1) followed immediately by the UTF-8 bytes of record n's JSON. h(0) is 64 zeros. Record 1 is the genesis record. Its JSON has exactly four keys: "adr_sha256", the SHA-256 of this document as defined in rule 14, in lowercase hexadecimal; "effective", the string "2026-10-01T07:00:00Z"; "rehearsal_head", the rehearsal chain head defined in rule 13, in lowercase hexadecimal; and "type", the string "genesis". The genesis hash is h(1). Records are only ever appended. 5. **No hand decisions.** No person registers, edits, resolves, withdraws or scores a call. Forecasting agents propose calls and set probabilities, the Resolver reaches verdicts, and deterministic code computes scores. These rules change only through a new ADR with its own hash, and a new ADR applies only to calls registered after it is published. 6. **Updates.** For an open call, SOTO may publish a new live probability with the time, the old and new values, a reason and the evidence. Updates are chained and never scored. Only p is scored. 7. **Resolution.** Within 72 hours after the deadline, the Resolver collects what the named sources published and runs 20 independent readings. A verdict of YES or NO needs at least 19 in agreement. Otherwise, the Resolver runs again seven days later, with everything published by then about events up to the deadline. If 19 of 20 still do not agree, the call is unresolvable. The verdict, the readings and the evidence links are chained. 8. **Scoring.** For a probability q, score(q) is ln(q) if the verdict is YES and ln(1 - q) if it is NO. A call's result is score(p) minus score(b). A positive result marks the call RESOLVED and a negative result marks it MISSED. An unresolvable or withdrawn call takes the lower of its two possible results and is marked MISSED. A call with a result under this rule is a scored call. 9. **Withdrawal.** Software withdraws an open call only if a named source stops existing, or if a repeat of the ambiguity test fails before the deadline. The reason is chained and the call is scored under rule 8. Nothing is deleted. 10. **Reporting.** Every published score shows the number of scored calls behind it. A calibration chart is published from 30 calls with a verdict of YES or NO. SOTO does not claim to beat its base rates until at least 50 scored calls have REFERENCE CLASS base rates and the sum of their results under rule 8 is above zero. Calls with JUDGMENT base rates do not count toward the 50 and are never used for that claim. 11. **Anchoring.** Every day at 23:59 Pacific time, starting on the day of the first registered call, software publishes the latest chain hash outside SOTO's own systems, as an OpenTimestamps proof and in a public post. Both are recorded in the chain. 12. **Corrections.** Errors in SOTO's coverage are corrected with a dated correction, chained and shown next to the original. Nothing is changed silently. 13. **Rehearsal.** The 19 calls SOTO published on September 29 and 30, 2026, numbered S1 to S12 and S17 to S23, are a rehearsal. They are archived and not scored, and their outcomes are published, unscored, as they become known. Four of them, S2, S8, S9 and S10, were withdrawn and stay on the record. No call was published as S13 to S16. The rehearsal chain has one record per call, in ascending order of number, S1 first. Each record is canonical JSON, serialized as in rule 4, with exactly these ten keys: "label", the call's number, such as "S1"; "claim"; "rule", the resolution rule text as registered; "deadline", a date as YYYY-MM-DD; "p", the registered probability as a whole number, or null; "b", the base rate as a whole number, or null; "published_at", in UTC as YYYY-MM-DDTHH:MM:SSZ; "status", which is "registered" or "withdrawn"; "withdrawn_at", in UTC in the same format, or null; and "withdrawal_reason", text or null. Hashes follow rule 4 with h(0) = 64 zeros and records numbered from 1. The rehearsal chain head is h(19), computed on October 1, 2026, from the records as they stood at the end of September 30, 2026, Pacific time. The 19 serialized records, one per line, are published unchanged in a file named rehearsal-chain.jsonl next to this document; anyone holding them can recompute the head. The head is 7286547c5f05e271b993008d2e7e47ddddb07751cf7b309b6cd48abac45084d8. 14. **Versions.** This document is identified by its SHA-256 hash, computed over its UTF-8 bytes from the start of the line that begins with "# ADR-000:" through its final newline. It is never edited. Changes are published as ADR-001 and onward.
Calls
C1: Will Mistral AI name a new enterprise customer or partner by Mar 30, 2027?
Claim. On or before 2027-03-30, Mistral AI will publicly announce a new named enterprise customer or partner that meets at least one disclosed scale criterion (1,000+ employees, a national/public-sector entity, a deal value of $10M USD or more, or a multi-year/multi-country deployment agreement).
Resolution rule. Resolves YES if, between 2026-10-01 and 2027-03-30T23:59:00-07:00, Mistral AI's official blog/press page (mistral.ai) or a named outlet (Reuters, Bloomberg, Financial Times, CNBC, TechCrunch, The Information) publishes a dated report naming a new enterprise customer or partner organization not already disclosed before 2026-10-01 (excludes Samsung, ESA, Morocco government, and the Munich hub partners), where the entity or deal meets at least one of: 1,000+ global employees, national/public-sector status, disclosed contract value of $10M USD or more, or a stated multi-year/multi-country deployment. The report must name the organization explicitly and confirm an active commercial or institutional relationship with Mistral. Anything else, including unconfirmed rumors or anonymized references, resolves NO.
Estimate. SOTO 85%; the obvious 82%. Deadline. 2027-03-30T23:59:00-07:00. Registered. Oct 1, 2026, 13:44 PT.
Sources that settle it. mistral.ai/news; Reuters (reuters.com); Bloomberg (bloomberg.com); Financial Times (ft.com); CNBC (cnbc.com); TechCrunch (techcrunch.com); The Information (theinformation.com)
C2: Will ElevenLabs state that Eleven v4 stays at $22 per million characters after Oct 12?
Claim. By Oct 16, 2026, ElevenLabs will have publicly stated, on or before that date, that the Eleven v4 API remains priced at $22 per million characters after Oct 12, 2026.
Resolution rule. Resolves YES only if, on or before 2026-10-16T23:59:00-07:00, ElevenLabs' official pricing page, its official blog, or its verified X account explicitly states that standard Eleven v4 API pricing is $22 per million characters effective after Oct 12, 2026. The statement must reference the standard self-serve API rate, not enterprise, custom, or negotiated pricing, and must not describe a new model, renamed tier, or different per-character rate. Silence, inferred continuity, archived pricing pages without a dated confirmation, or third-party reporting does not qualify. Absence of any such statement by the deadline resolves NO.
Estimate. SOTO 12%; the obvious 35%. Deadline. 2026-10-16T23:59:00-07:00. Registered. Oct 2, 2026, 11:07 PT.
Sources that settle it. https://elevenlabs.io/pricing; https://elevenlabs.io/blog; https://x.com/elevenlabsio
C3: Will OpenAI publish a price for extra Dots usage by Oct 31?
Claim. By October 31, 2026, OpenAI publicly publishes a specific price (dollar amount, rate, or credit cost) for additional Dots usage or for Dots-related compute beyond what existing ChatGPT subscription plans include.
Resolution rule. Resolves YES if, on or before 2026-10-31T23:59:00-07:00, OpenAI's own website, pricing page, or help center publishes a specific numeric price, rate, or credit cost for extra Dots usage or Dots compute beyond stated plan allowances, OR a tier-1 outlet (e.g., The Verge, NYT, WSJ, Reuters, Bloomberg, TechCrunch) publishes an article quoting or citing that OpenAI-sourced price. The statement must include a concrete figure (e.g., dollars, credits, or per-unit rate), not a vague reference to 'future pricing' or 'paid tiers.' Anything else, including statements made before October 2, 2026, resolves NO.
Estimate. SOTO 6%; the obvious 10%. Deadline. 2026-10-31T23:59:00-07:00. Registered. Oct 2, 2026, 11:07 PT.
Sources that settle it. openai.com/pricing; help.openai.com (OpenAI Help Center); theverge.com; nytimes.com; wsj.com; reuters.com; bloomberg.com; techcrunch.com
C4: Will a model that runs offline on a laptop or phone match the top Artificial Analysis score of Oct 2, 2026 by Apr 2, 2027?
Claim. On or before 2027-04-02, a model that runs fully offline on one consumer laptop or phone (up to 64 GB of memory, no network use at inference) scores at or above 58 on the Artificial Analysis Intelligence Index v4.3, the score of the #1 model on that index on 2026-10-02 (Claude Opus 5.5, max with fallback).
Resolution rule. Resolves YES if, on or before 2027-04-02 23:59 Pacific Time, Artificial Analysis (artificialanalysis.ai) publishes an Intelligence Index score of 58 or higher (on index v4.3, or on a later version where the 2026-10-02 leader, Claude Opus 5.5 max with fallback, is re-scored: then the leader's re-scored value) for a model whose weights are publicly downloadable and which has a documented run fully offline on one consumer laptop or phone with up to 64 GB of memory. The documented run may be by Artificial Analysis or a public, reproducible run by anyone. Resolves NO otherwise. Not counted: benchmark results published only by the model's maker, models offered only through an API, quantized variants that Artificial Analysis has not scored, and any index other than Artificial Analysis. The 2026-10-02 reference page is archived at https://web.archive.org/web/20261003030249/https://artificialanalysis.ai/leaderboards/models.
Estimate. SOTO 8%; the obvious 35%. Deadline. 2027-04-02T23:59:00-07:00. Registered. Oct 3, 2026, 00:17 PT.
Sources that settle it. https://artificialanalysis.ai/leaderboards/models; https://web.archive.org/web/20261003030249/https://artificialanalysis.ai/leaderboards/models; https://conway.tech/