System Atlas
Public architecture concepts
An illustrative map of research and review boundaries. No private roster, operational census, runtime schedule, account data, or deployment status is published.
Wider than the screen.Scroll the map sideways, or put focus on it and use the arrow keys.
Multi-Model Prediction Market Intelligence
What the system is allowed to read, where inference is bought, and the limits on anything it sends back out.
Market data arrives through one path per vertical, and a feed that cannot prove its own freshness is refused rather than guessed at.
- Venue adapters behind explicit access policy
- Crypto research from closed-candle evidence
- Options-chain provenance and freshness checks
- Keyless FX rates plus intraday candles
- Freshness derived server-side
- Provider routing with explicit fallback
- Monthly cap per provider
- Cost-aware provider routing
- Never fabricate: unmeasured, no call
- Cited research, cached to memory
- Loopback or constant-time token
- Typed confirm on destructive routes
- Digest-pinned images, loopback ports
- Secrets never enter version control
- Cash derived, never stored
- Hard brake stops new stakes
- Circuit breaker resets cash only
- Exposure ceiling checked each pass
The panel that votes, the rubric it answers, the gate that decides which plays are worth naming, and the ladder that governs a seat's tenure.
Dozens of model seats each form their own opinion on every market, and influence is earned from graded results rather than from brand.
- Independent analytical roles
- Bayesian-shrunk win-loss ranking
- Provider-source tag on every seat
- Provider routing with explicit fallback
- Unanswered votes marked and excluded
- Parallel fan-out per market
- A structured set of factors, weighted and sport-tiered
- Per-sport tiers: tennis, MLB, football
- Criteria reachability audited
- Crypto FLAT is explicit abstention
- Badge gate: 5 voters, 90% agreement
- No rate without its null
- Null from entry price, never 0.50
- Current-epoch scoping only
- Cross-row golden board
- Leaderboard by shrunk win rate
- Reviewed roster changes with explicit authorization
- Operator ban registry
- Chronic-loser reset, top-skill re-seed
- Parked seats leave every money query
- Vector indexes for similarity retrieval
- Separate memory collections by purpose
- Typed marker ledger, never erased
- Additive migrations with bounded lock waits
Redis · forex store
Scan loop · settlement pass
Orchestrator row
CEO-gated exception
The one real-money path stays behind a person
Three things a seat can hold, and the codebase keeps them apart. A role is given. A lens is available to every seat. A record is earned, and it is never transferable.
The role examples describe analytical perspectives, not current model assignments.
- Reasoning roles from calibration to counter-case
- The role is written into that seat's own prompt
- A missing role is a plain default, never an error
- Nine named donors may lend a method
- Method only: no pick, no confidence, no wording
- The borrower keeps its own native role
- Fixed order: native first, borrowed last
- Core and derived analytical methods
- Crypto and options research methods
- FX strategy definitions and implementation states
- Four matrix states, never collapsed into one
- Core is a seat's own thinking, and it is protected
- Only an acquired habit can ever be stripped
- Unknown provenance now reads as protected
- Protected keys never reach the reviewer at all
- Three kinds refuse transfer: skill, calibration, heuristic
- A graft may move method, never a record
- The rule raises rather than warns
- An unreadable ban list denies, never allows
- Versioned playbooks with explicit retirement
- Integration bundles with explicit permissions
- One design contract every screen obeys
- A merge checklist that refuses unread work
Where a claim becomes a fact and a fact is allowed to change behaviour. Every instrument here reads settled outcomes, and the ones that could act on their own are off by default.
The boundary between an opinion and a fact. Before it, votes; after it, graded evidence, and only graded evidence is allowed to teach the system anything.
- Atomic settlement, two independent resolvers
- Three-channel synthetic gate on every read
- Epoch markers, history never erased
- A placeholder vote can never become evidence
- Not measured, measured-none and measured k of N stay apart
- Evidence-gated skill development
- Validated cells mint as warming, never active
- Calibration health: dead-signal check
- Criteria: declared, reachable, scored
- Explicit execution and availability states
- Evidence-based review per role and domain
- Worst styles by shrunk ranking
- Core styles suppressed, never deleted
- Opt-in review with workload limits
- Write boundary refuses core bans
- Eight losses in a row on one sport
- Recovery marker opens a new epoch
- Bad-skill strip, core preserved
- Method graft from the sport leader
- Borrowed prior, additive prompt overlay
- Three evidence states, never merged
- Beta posterior plus Wilson bound
- Trailing loss streak against a threshold
- Health and reachability blocks
- Rule table: reset, abstain, size down
- Typed inputs checked at source
- Unregistered identities are loud
- Known-gap ratchet, stale gaps go red
- A cognition-engine wire per seat
The ledger is the boundary. Before it, opinions. After it, evidence, and nothing on this page is allowed to learn from anything that has not crossed it.
This example distinguishes evaluation from activation. No private stage census, enabled-state inventory, or operating schedule is published.
The four books the system actually trades, the screens an operator reads them on, and the shared design language those screens are held to.
Separate research domains retain their own evidence and evaluation records.
- PolyPredictions: tennis, basketball, MLB, football
- PolyCrypto and Flips: closed-candle rows
- PolyOptions: research live, trading dormant
- PolyForex: paper, fail-closed on freshness
- Sports rows: 17 to 19 tabs into 6 groups
- Observation routes and explicit availability states
- Orchestrator row: cost, memory, agents
- Loopback dashboard, eleven data states
- Shared explorer contract, per-sport adapters
- Outcome scrubber, Wilson sample map
- Complete seat census, explicit scopes
- Deferred states: loading, ready, error
- Readiness verified by browser render
- Lattice, pulse ring, streak bar
- Four states by shape, not by hue
- DOM-built, no random values
- Plain-SVG fallback when absent
- Graphic contrast floor 3:1, measured
How the thing stays up, and how it gets written. Many AI engineering sessions edit this codebase at once, so the coordination rules are part of the architecture.
Health is measured as the age of what a job last wrote, never as the existence of its process. A live process serving stale work is the failure this is built to catch.
- Healthchecks probe liveness, not process ids
- Restart policy lives in the base file
- Restart acts on exit, not on health
- Test containers are a memory tenant
- Stop and start, never force-recreate
- Settlement monitoring with explicit stale states
- Throttled skill review
- Operational observability and health reporting
- Many AI coding sessions, concurrently
- Append-only coordination log with expiry
- Incident registry tied to guard tests
- Playbooks govern how work is built and merged
- Sector-shell design contract
- Integration contracts and reviewable capabilities
- One retired playbook kept for history
- Operational observability
- Trading-agent metrics scraped
- Always-firing dead-man alert
- Synthetic browser-path probe
Several AI engineering sessions edit this codebase at the same time, so the coordination ledger, the restart discipline and the shared design contract are architecture, not notes beside it.
A landed change is not a live change until the running process matches it, which is why every claim on the operations side is verified by rendering the real page rather than by reading a green status.
This table describes responsibilities and configurable scheduling policies. Private runtime intervals and current operating status are not published.
Synthetic · no live readingsKeeping not measured, measured and none found, and measured k of N apart is a rule the illustrative architecture enforces too. A missing state must never fall through to the most reassuring one.
REPORTING / ARTIFACTS
Dashboards · Docs · Audits & Operational Records
Components and their design relationships
The map above is the executive view. This is the illustrative architecture: illustrative components across nine categories, from data feeds and the model panel to memory, execution, the safety stack, the engineering roles, integrations, and the human interfaces. Each row is one category. The lines show conceptual component relationships, kept faint so the structure reads first.
Wider than the screen.Scroll the map sideways, or put focus on it and use the arrow keys.
Why dozens of models vote instead of one model deciding
This section illustrates component responsibilities and evidence boundaries. Private inventories, thresholds, and activation settings are intentionally omitted.
| Mechanism | How it works | Why it was built | Failure it prevents |
|---|---|---|---|
| Measured skill | MEASURED, NOT ASSUMED Every seat carries a resolved win-loss rate with its confidence interval, its calibration gap (claimed confidence versus realized accuracy) and its record against the price the market already implied. That measurement is illustrative architecture; feeding it back into how the panel trades stays gated until the evidence clears its floors. |
Model quality varies by sport, market type and time, and a rate without a sample size or a baseline flatters itself. | A once-good model quietly dragging every decision down, and a lucky streak being read as skill. |
| Criteria rubric | STRUCTURED PROMPTING Structured factors, carried consistently by the factor list and the weight table, tiered per sport, and every seat must commit to a pick. Confidence is the weighted average of per-factor scores, not a vibe, and an audit proves each key is both prompted and weighted. |
Cross-model confidence is only comparable when every model scores the same factors the same way. | Unauditable gut-feel confidence, and a criterion that is asked for but can never be scored. |
| Golden play contract | A RATE NEEDS ITS NULL One shared definition across every row: a badge needs at least 5 distinct voters, 90 percent agreement and a licence of 20 settled units in that band, and no win rate is published without the price-implied baseline beside it. |
Consensus strength should be visible and priced, and a rate quoted without the market's own expectation is not evidence of edge. | Treating a lone opinion and a full-panel lock the same way, and publishing a flattering rate with no baseline. |
| Seat governance | PARK · BAN · RESET A seat can be parked out of consensus and out of every money query, banned permanently by the operator, retired from inference with its history kept, or reset onto a fresh epoch. A policy-defined review trigger ejects; automatic roster replacement stays switched off by default. |
An ensemble is only as honest as its membership rules. Seats are earned, never tenured, and a reset must never erase the record that justified it. | Zombie seats surviving on reputation, and a recovery that destroys its own evidence. |
| Research route | CITED LIVE RESEARCH A dedicated research path pulls citation-backed live web intelligence, with an orchestrator fallback, and caches findings into vector memory for reuse. |
Markets move on news. Evidence must arrive with citations, timestamps, and a memory trail. | Stale narratives being priced as fresh information. |
How each model turns a market into a number: the sport-tiered criteria rubric
Every model on the panel scores the same market against the same factors, in the same order, on the same scale. That is what makes dozens of different models comparable: their confidence is not a gut feeling, it is a weighted average of explicit judgments. The twelve below are the universal core that every sport shares, drawn from A structured set of factors; tennis, baseball and football each add their own tier on top. Each factor is scored from 0 to 1 toward the model's own pick, where 0.5 means unknown and is dropped so it cannot dilute the signal. Factors are weighted by how reliably they predict, so a market-edge read counts far more than a weather note.
| Factor | What it evaluates | Weight |
|---|---|---|
| Market edge | The model's fair-value estimate versus the market-implied price. The strongest reality-check on raw opinion. | 0.15 |
| Injuries & availability | Who is out, questionable, or returning; roster availability at tip-off. | 0.12 |
| Kelly sizing | An edge-magnitude check: how large a position the estimated edge actually justifies. | 0.12 |
| Record & form | Win/loss record and results across recent games. | 0.10 |
| Momentum | Hot or cold trend beyond raw record; win streaks. | 0.09 |
| Head-to-head | Prior matchups between the two sides. | 0.08 |
| Calibration vs history | The model's own historical accuracy on similar markets. Its track record grades its own confidence. | 0.08 |
| Home / away | Venue and home-field advantage. | 0.07 |
| Rest & schedule | Back-to-backs, travel, and schedule congestion. | 0.06 |
| Sentiment & news | Public sentiment, quotes, morale, and line-movement signal. | 0.05 |
| Age & peak window | Career stage and physical peak. | 0.05 |
| Weather & conditions | Wind, rain, temperature, and venue conditions. | 0.03 |
The factor scores combine into a single quant confidence by weighted average. Factors scored 0.5 (unknown) are excluded from both sides of the ratio, so missing information lowers certainty instead of faking it.
Final confidence is 0.6 times the quant score plus 0.4 times the model's own stated confidence. The rubric keeps models honest; the blend keeps their genuine read in the mix.
A tiny edge is forced to match the market near 0.60; a larger edge earns more; an oversized claimed edge is penalized because it historically underperforms. Weak-history leagues carry hard confidence caps.
This section illustrates component responsibilities and evidence boundaries. Private inventories, thresholds, and activation settings are intentionally omitted.
This section illustrates component responsibilities and evidence boundaries. Private inventories, thresholds, and activation settings are intentionally omitted.
Once a market resolves, the calibration factor and the panel weights update from what actually happened, so the same factors are scored a little more wisely next time.
The learning loop: every prediction becomes training signal
This is an architectural example. Current operating status, runtime measurements, and private configuration are not verified or published here.
Monitors resolve each market (API-first, model fallback), stamp the prediction with its result, and an accountability task backfills per-model accuracy records and the calibration log.
Claimed confidence is compared with realized win rate per seat and per row, and a health check catches a calibration curve that has collapsed to a constant instead of reporting it as a fit.
This is an architectural example. Current operating status, runtime measurements, and private configuration are not verified or published here.
Outcomes become embeddings: similar-situation recall informs the next pick, error-to-fix memory suggests repairs on repeat failures, and training scaffolds from model meetings are prepended to future prompts.
Memory architecture: vector retrieval and a marker ledger
Similarity memory and durable event history answer different questions. Storage dimensions, table inventories, and private deployment details are omitted.
Operator trades, unusual plays, per-model accuracy, chat Q&A, error-to-fix pairs, game outcomes, news sentiment, model-meeting insights, market text, and lifecycle evidence. Each table answers one recall question well.
High-dimension embeddings exceed the standard HNSW limit, so vectors are indexed as half-precision with cosine ops. Approximate search trades recall, latency, and storage; validate the choice on representative queries.
Every reset, graft, sweep and bundle writes one typed, additive row. Vectors answer "what is similar"; the marker ledger answers "what happened to this seat, and when." Nothing is ever mutated, so a reset is a time floor rather than a deletion.
In-memory cache, cross-restart alert dedup, and signal pub/sub. Notifications survive restarts without double-firing, and hot data never touches the database.
Git history alone does not back up a database, secrets or runtime state, so backup and restore are documented separately and audited by a read-only readiness check. Recovery is a written runbook, not an assumption.
Recall is filtered by distance thresholds and hit counters, cached answers are reused when close enough, and repeated model discussions are deduplicated before they waste tokens.
The engineering fleet: many concurrent AI sessions coordinated through a work ledger
Engineering changes benefit from isolated workspaces, explicit ownership, independent review, and reproducible checks.
Engineering changes benefit from isolated workspaces, explicit ownership, independent review, and reproducible checks.
Engineering changes benefit from isolated workspaces, explicit ownership, independent review, and reproducible checks.
Engineering changes benefit from isolated workspaces, explicit ownership, independent review, and reproducible checks.
Each significant incident is written down with the exact test that now enforces the fix, so a class of bug cannot quietly return. Automated remediation stays recommendation-only: an agent proposes, a human decides.
Reviewable skills and integration boundaries
This section illustrates component responsibilities and evidence boundaries. Private inventories, thresholds, and activation settings are intentionally omitted.
Preview snapshots, edge-config flags, dashboard deploys, and a release guard so the public surface ships through checks, never by hand.
Shadow backtests, dataset curation, space evaluation runs, and a scheduled model scout feeding candidates to the bench.
DFS projections, injury pulse, lineup watch, and reanalysis triggers when ground truth changes under an open position.
Drawdown circuit breaker, bankroll rebalancing, settlement idleness monitors, and watchdog operations for the money path.
Early-edge alerts, confidence calibration, edge attribution, and a postgame loop that turns every result into analysis.
Skill-coverage audits, plugin readiness orchestration, page release readiness, and a scheduled go/no-go verdict.
Roster sync, research databases, and project hubs in Notion; research memory, loss-cluster analysis, and postmortem mining in Dovetail.
Issue sync and SLO burn tracking, email briefs with delivery guards, and a chart CDN pipeline with visual audits.
Warehouse queries, container deploys, storage artifacts, pub/sub bridging, managed inference, and secret rotation recipes.
Risk, safety & release discipline: engineered like it can fail
An autonomous system touching real markets gets the full defensive stack: position sizing with hard caps, multi-level kill switches, strict separation between simulated and real money, injection-scrubbed prompts, masked errors, authenticated surfaces, and a public boundary designed so leaking is structurally difficult.
Fractional Kelly sizing with a hard per-position cap, confidence-tier stake ladders, and minimum edge and expected-value floors before any stake is placed.
Automatic staking halts at the model, league, and model-league level when resolved win rates or P&L fall below floors. Plus a pause flag and a phone-side switch that stops the whole stack.
Models trade virtual per-model bankrolls. Real accounts are tracked read-only, live actions require operator approval, and the browser execution bridge ships hard-capped and disabled by default.
The design calls for constrained inputs, protected secrets, and reviewed publication boundaries. These mechanisms reduce risk; they do not establish complete prevention.
The design calls for constrained inputs, protected secrets, and reviewed publication boundaries. These mechanisms reduce risk; they do not establish complete prevention.
Dashboards bind to localhost by default; remote access requires a token compared in constant time; write endpoints are gated and rate-limited.
The design calls for constrained inputs, protected secrets, and reviewed publication boundaries. These mechanisms reduce risk; they do not establish complete prevention.
Backup-readiness audits, fresh-clone recovery runbooks, machine-migration parity reports, and an append-only sync manifest make the whole system rebuildable from documentation.
One market, end to end
An illustrative path from a research question to a reviewed outcome.
Markets stream in from the prediction venue through a whitelisted, deduplicated API client and land in Postgres with full provenance. Each of the four verticals has its own door, and every one of them fails closed on stale data.
Sharp sportsbook lines, injury and lineup intelligence, live scores, websocket prices, and cited web research attach context to each market.
The vector brain surfaces similar past situations and their outcomes; the marker ledger contributes what has already happened to this seat, including any reset or graft.
The full panel runs the sport-tiered criteria rubric in parallel and each seat commits to a pick. A seat that does not answer is marked synthetic and excluded from every money query.
The golden-play contract decides whether this is a badge-worthy play: enough distinct voters, enough agreement, a licensed sample, and a published rate only ever beside the price-implied baseline.
Risk controls apply: confidence, edge, and expected-value floors, fractional Kelly sizing, stake ladders, hard brakes, and operator approval for anything real.
Monitors track every open position, resolve outcomes API-first with model fallback, and stamp each prediction with its result.
Calibration updates, rankings move, embeddings absorb the outcome, and the substrate census records which learning stage actually ran. Where the evidence clears its floor, a losing seat is reset and re-taught the leading seat's method.