POLYMIND

Multi-model research, evidence, and accountable decisions.

Explore the design of a multi-model research platform: independent estimates, evidence gates, measured uncertainty, and human approval. This conceptual showcase does not publish private runtime configuration or claim live performance.

Design Rationale

System Atlas

Public architecture concepts

An illustrative map of research and review boundaries. No private roster, operational census, runtime schedule, account data, or deployment status is published.

Wider than the screen.Scroll the map sideways, or put focus on it and use the arrow keys.

POLYMIND

Multi-Model Prediction Market Intelligence

Public Showcase · Synthetic Metrics
Trace a path Click any module for what it does, why it was built this way, and what it connects to
01 Intake and perimeter

What the system is allowed to read, where inference is bought, and the limits on anything it sends back out.

MARKET DATA INTAKEFour verticals · fail-closed feeds

Market data arrives through one path per vertical, and a feed that cannot prove its own freshness is refused rather than guessed at.

  • Venue adapters behind explicit access policy
  • Crypto research from closed-candle evidence
  • Options-chain provenance and freshness checks
  • Keyless FX rates plus intraday candles
  • Freshness derived server-side
PROVIDER ROUTESInference lanes · research feed
  • Provider routing with explicit fallback
  • Monthly cap per provider
  • Cost-aware provider routing
  • Never fabricate: unmeasured, no call
  • Cited research, cached to memory
PUBLIC BOUNDARYFail-closed release discipline
  • Loopback or constant-time token
  • Typed confirm on destructive routes
  • Digest-pinned images, loopback ports
  • Secrets never enter version control
RISK CONTROLSBrake · breaker · floor
  • Cash derived, never stored
  • Hard brake stops new stakes
  • Circuit breaker resets cash only
  • Exposure ceiling checked each pass
02 Model cognition

The panel that votes, the rubric it answers, the gate that decides which plays are worth naming, and the ladder that governs a seat's tenure.

MODEL VOTING PANELillustrative architecture · measured skill

Dozens of model seats each form their own opinion on every market, and influence is earned from graded results rather than from brand.

  • Independent analytical roles
  • Bayesian-shrunk win-loss ranking
  • Provider-source tag on every seat
  • Provider routing with explicit fallback
  • Unanswered votes marked and excluded
DEEP-QUANT ANALYSTSport-scoped criteria rubric
  • Parallel fan-out per market
  • A structured set of factors, weighted and sport-tiered
  • Per-sport tiers: tennis, MLB, football
  • Criteria reachability audited
  • Crypto FLAT is explicit abstention
SIGNAL ENGINEGolden plays · null-adjusted
  • Badge gate: 5 voters, 90% agreement
  • No rate without its null
  • Null from entry price, never 0.50
  • Current-epoch scoping only
  • Cross-row golden board
MODEL LIFECYCLEPark · ban · reset ladder
  • Leaderboard by shrunk win rate
  • Reviewed roster changes with explicit authorization
  • Operator ban registry
  • Chronic-loser reset, top-skill re-seed
  • Parked seats leave every money query
VECTOR BRAINVector memory collections · marker ledger
  • Vector indexes for similarity retrieval
  • Separate memory collections by purpose
  • Typed marker ledger, never erased
  • Additive migrations with bounded lock waits
Knowledge LayerPostgres · pgvector
Redis · forex store
POLYMIND Orchestration Core

Scan loop · settlement pass
Orchestrator row

Execution LayerPaper only
CEO-gated exception
CEO Approval GatesHuman sign-off on every real action
The one real-money path stays behind a person
03 Role · Lens · Record

Three things a seat can hold, and the codebase keeps them apart. A role is given. A lens is available to every seat. A record is earned, and it is never transferable.

Research rolesRole · one named job per seat
Skeptical assumption auditor Web-grounded source validator Algorithmic factor debugger Long-context evidence planner Bayesian market-structure analyst

The role examples describe analytical perspectives, not current model assignments.

  • Reasoning roles from calibration to counter-case
  • The role is written into that seat's own prompt
  • A missing role is a plain default, never an error
DONOR DISCIPLINE OVERLAYRole · additive, never a replacement
  • Nine named donors may lend a method
  • Method only: no pick, no confidence, no wording
  • The borrower keeps its own native role
  • Fixed order: native first, borrowed last
STYLE LIBRARIESLens · shipped to all, assigned to none
  • Core and derived analytical methods
  • Crypto and options research methods
  • FX strategy definitions and implementation states
  • Four matrix states, never collapsed into one
CORE VERSUS ACQUIREDRecord · four tests for original thinking
  • Core is a seat's own thinking, and it is protected
  • Only an acquired habit can ever be stripped
  • Unknown provenance now reads as protected
  • Protected keys never reach the reviewer at all
NEVER TRANSFERABLERecord · it stays where it was earned
  • Three kinds refuse transfer: skill, calibration, heuristic
  • A graft may move method, never a record
  • The rule raises rather than warns
  • An unreadable ban list denies, never allows
AGENT PLAYBOOKSRole · the contract the builders work to
  • Versioned playbooks with explicit retirement
  • Integration bundles with explicit permissions
  • One design contract every screen obeys
  • A merge checklist that refuses unread work
Open the detail
04 Evidence and self-learning

Where a claim becomes a fact and a fact is allowed to change behaviour. Every instrument here reads settled outcomes, and the ones that could act on their own are off by default.

OUTCOME LEDGERSettlement · evidence gate · epochs

The boundary between an opinion and a fact. Before it, votes; after it, graded evidence, and only graded evidence is allowed to teach the system anything.

  • Atomic settlement, two independent resolvers
  • Three-channel synthetic gate on every read
  • Epoch markers, history never erased
  • A placeholder vote can never become evidence
  • Not measured, measured-none and measured k of N stay apart
SELF-LEARNING SUBSTRATESkill pass · health · census
  • Evidence-gated skill development
  • Validated cells mint as warming, never active
  • Calibration health: dead-signal check
  • Criteria: declared, reachable, scored
  • Explicit execution and availability states
BAD-SKILL RECOGNIZERCore-skill partition · sub-scopes
  • Evidence-based review per role and domain
  • Worst styles by shrunk ranking
  • Core styles suppressed, never deleted
  • Opt-in review with workload limits
  • Write boundary refuses core bans
HOMEOSTATIC RESET LOOPEight losses · graft method
  • Eight losses in a row on one sport
  • Recovery marker opens a new epoch
  • Bad-skill strip, core preserved
  • Method graft from the sport leader
  • Borrowed prior, additive prompt overlay
NAXIKILL COGNITIONEvidence to next action, no LLM
  • Three evidence states, never merged
  • Beta posterior plus Wilson bound
  • Trailing loss streak against a threshold
  • Health and reachability blocks
  • Rule table: reset, abstain, size down
PIPELINE REGISTRYEvery seat declares every wire
  • Typed inputs checked at source
  • Unregistered identities are loud
  • Known-gap ratchet, stale gaps go red
  • A cognition-engine wire per seat
How this band reads

The ledger is the boundary. Before it, opinions. After it, evidence, and nothing on this page is allowed to learn from anything that has not crossed it.

This example distinguishes evaluation from activation. No private stage census, enabled-state inventory, or operating schedule is published.

05 Rows and surfaces

The four books the system actually trades, the screens an operator reads them on, and the shared design language those screens are held to.

TRADING ROWSFour verticals · sports sub-scopes

Separate research domains retain their own evidence and evaluation records.

  • PolyPredictions: tennis, basketball, MLB, football
  • PolyCrypto and Flips: closed-candle rows
  • PolyOptions: research live, trading dormant
  • PolyForex: paper, fail-closed on freshness
  • Sports rows: 17 to 19 tabs into 6 groups
COMMAND SURFACESSector shell · rows · panels
  • Observation routes and explicit availability states
  • Orchestrator row: cost, memory, agents
  • Loopback dashboard, eleven data states
EVIDENCE EXPLORERSInteractive record explorers
  • Shared explorer contract, per-sport adapters
  • Outcome scrubber, Wilson sample map
  • Complete seat census, explicit scopes
  • Deferred states: loading, ready, error
  • Readiness verified by browser render
PLASMA NEURON KITShared cognition design language
  • Lattice, pulse ring, streak bar
  • Four states by shape, not by hue
  • DOM-built, no random values
  • Plain-SVG fallback when absent
  • Graphic contrast floor 3:1, measured
06 Operations and the build

How the thing stays up, and how it gets written. Many AI engineering sessions edit this codebase at once, so the coordination rules are part of the architecture.

OPERATIONSOperational responsibilities

Health is measured as the age of what a job last wrote, never as the existence of its process. A live process serving stale work is the failure this is built to catch.

  • Healthchecks probe liveness, not process ids
  • Restart policy lives in the base file
  • Restart acts on exit, not on health
  • Test containers are a memory tenant
  • Stop and start, never force-recreate
SCHEDULED STAGESOne watchdog loop · inline stages
  • Settlement monitoring with explicit stale states
  • Throttled skill review
  • Operational observability and health reporting
MULTI-AGENT FLEETLedger-coordinated AI sessions
  • Many AI coding sessions, concurrently
  • Append-only coordination log with expiry
  • Incident registry tied to guard tests
SKILLS & AGENT CONTRACTSReviewed capabilities · explicit contracts
  • Playbooks govern how work is built and merged
  • Sector-shell design contract
  • Integration contracts and reviewable capabilities
  • One retired playbook kept for history
OBSERVABILITYMetrics · dashboards · health probe
  • Operational observability
  • Trading-agent metrics scraped
  • Always-firing dead-man alert
  • Synthetic browser-path probe
Why the build is on the map

Several AI engineering sessions edit this codebase at the same time, so the coordination ledger, the restart discipline and the shared design contract are architecture, not notes beside it.

A landed change is not a live change until the running process matches it, which is why every claim on the operations side is verified by rendering the real page rather than by reading a green status.

CadenceScheduling responsibilities

This table describes responsibilities and configurable scheduling policies. Private runtime intervals and current operating status are not published.

Synthetic · no live readings
LoopScheduling policyReading on this page

Keeping not measured, measured and none found, and measured k of N apart is a rule the illustrative architecture enforces too. A missing state must never fall through to the most reassuring one.

REPORTING / ARTIFACTS

Dashboards · Docs · Audits & Operational Records

Operations Dashboard
Row Readiness Docs
Critical Fixes Registry
Security Audit Reports
Calibration & Skill Tabs
Backup & Restore Reference
Browser Readiness Probes
Engineering Work Log
Project Memory Log
Public-Style Dashboard
Public architecture showcase · Synthetic metrics only · No credentials or trading records
System Atlas · Expanded

Components and their design relationships

The map above is the executive view. This is the illustrative architecture: illustrative components across nine categories, from data feeds and the model panel to memory, execution, the safety stack, the engineering roles, integrations, and the human interfaces. Each row is one category. The lines show conceptual component relationships, kept faint so the structure reads first.

Hover any node to light up only its connections and mute the rest. Click a node for what it does, why it exists, what powers it, and what it connects to.

Wider than the screen.Scroll the map sideways, or put focus on it and use the arrow keys.

POLYMIND SYSTEMS MAPTools · Skills · Systems in Conversation
Architecture component groups
Design Rationale · 01

Why dozens of models vote instead of one model deciding

This section illustrates component responsibilities and evidence boundaries. Private inventories, thresholds, and activation settings are intentionally omitted.

MechanismHow it worksWhy it was builtFailure it prevents
Measured skill MEASURED, NOT ASSUMED
Every seat carries a resolved win-loss rate with its confidence interval, its calibration gap (claimed confidence versus realized accuracy) and its record against the price the market already implied. That measurement is illustrative architecture; feeding it back into how the panel trades stays gated until the evidence clears its floors.
Model quality varies by sport, market type and time, and a rate without a sample size or a baseline flatters itself. A once-good model quietly dragging every decision down, and a lucky streak being read as skill.
Criteria rubric STRUCTURED PROMPTING
Structured factors, carried consistently by the factor list and the weight table, tiered per sport, and every seat must commit to a pick. Confidence is the weighted average of per-factor scores, not a vibe, and an audit proves each key is both prompted and weighted.
Cross-model confidence is only comparable when every model scores the same factors the same way. Unauditable gut-feel confidence, and a criterion that is asked for but can never be scored.
Golden play contract A RATE NEEDS ITS NULL
One shared definition across every row: a badge needs at least 5 distinct voters, 90 percent agreement and a licence of 20 settled units in that band, and no win rate is published without the price-implied baseline beside it.
Consensus strength should be visible and priced, and a rate quoted without the market's own expectation is not evidence of edge. Treating a lone opinion and a full-panel lock the same way, and publishing a flattering rate with no baseline.
Seat governance PARK · BAN · RESET
A seat can be parked out of consensus and out of every money query, banned permanently by the operator, retired from inference with its history kept, or reset onto a fresh epoch. A policy-defined review trigger ejects; automatic roster replacement stays switched off by default.
An ensemble is only as honest as its membership rules. Seats are earned, never tenured, and a reset must never erase the record that justified it. Zombie seats surviving on reputation, and a recovery that destroys its own evidence.
Research route CITED LIVE RESEARCH
A dedicated research path pulls citation-backed live web intelligence, with an orchestrator fallback, and caches findings into vector memory for reuse.
Markets move on news. Evidence must arrive with citations, timestamps, and a memory trail. Stale narratives being priced as fresh information.
Prediction Method · The Rubric

How each model turns a market into a number: the sport-tiered criteria rubric

Every model on the panel scores the same market against the same factors, in the same order, on the same scale. That is what makes dozens of different models comparable: their confidence is not a gut feeling, it is a weighted average of explicit judgments. The twelve below are the universal core that every sport shares, drawn from A structured set of factors; tennis, baseball and football each add their own tier on top. Each factor is scored from 0 to 1 toward the model's own pick, where 0.5 means unknown and is dropped so it cannot dilute the signal. Factors are weighted by how reliably they predict, so a market-edge read counts far more than a weather note.

FactorWhat it evaluatesWeight
Market edgeThe model's fair-value estimate versus the market-implied price. The strongest reality-check on raw opinion.
0.15
Injuries & availabilityWho is out, questionable, or returning; roster availability at tip-off.
0.12
Kelly sizingAn edge-magnitude check: how large a position the estimated edge actually justifies.
0.12
Record & formWin/loss record and results across recent games.
0.10
MomentumHot or cold trend beyond raw record; win streaks.
0.09
Head-to-headPrior matchups between the two sides.
0.08
Calibration vs historyThe model's own historical accuracy on similar markets. Its track record grades its own confidence.
0.08
Home / awayVenue and home-field advantage.
0.07
Rest & scheduleBack-to-backs, travel, and schedule congestion.
0.06
Sentiment & newsPublic sentiment, quotes, morale, and line-movement signal.
0.05
Age & peak windowCareer stage and physical peak.
0.05
Weather & conditionsWind, rain, temperature, and venue conditions.
0.03
Weighted averageHow the score forms

The factor scores combine into a single quant confidence by weighted average. Factors scored 0.5 (unknown) are excluded from both sides of the ratio, so missing information lowers certainty instead of faking it.

Sixty / forty blendDiscipline without robotics

Final confidence is 0.6 times the quant score plus 0.4 times the model's own stated confidence. The rubric keeps models honest; the blend keeps their genuine read in the mix.

Edge gates & capsOverconfidence brakes

A tiny edge is forced to match the market near 0.60; a larger edge earns more; an oversized claimed edge is penalized because it historically underperforms. Weak-history leagues carry hard confidence caps.

Sport overlaysDomain enrichment

This section illustrates component responsibilities and evidence boundaries. Private inventories, thresholds, and activation settings are intentionally omitted.

Mandatory pickNo abstaining

This section illustrates component responsibilities and evidence boundaries. Private inventories, thresholds, and activation settings are intentionally omitted.

Fed back, then re-weightedClosing the loop

Once a market resolves, the calibration factor and the panel weights update from what actually happened, so the same factors are scored a little more wisely next time.

Design Rationale · 02

The learning loop: every prediction becomes training signal

This is an architectural example. Current operating status, runtime measurements, and private configuration are not verified or published here.

Grade every outcome

Monitors resolve each market (API-first, model fallback), stamp the prediction with its result, and an accountability task backfills per-model accuracy records and the calibration log.

Score the calibration

Claimed confidence is compared with realized win rate per seat and per row, and a health check catches a calibration curve that has collapsed to a constant instead of reporting it as a fit.

Rank, and act where it is earned

This is an architectural example. Current operating status, runtime measurements, and private configuration are not verified or published here.

Feed memory forward

Outcomes become embeddings: similar-situation recall informs the next pick, error-to-fix memory suggests repairs on repeat failures, and training scaffolds from model meetings are prepended to future prompts.

Design Rationale · 03

Memory architecture: vector retrieval and a marker ledger

Similarity memory and durable event history answer different questions. Storage dimensions, table inventories, and private deployment details are omitted.

Vector MemoryWhat it remembers

Operator trades, unusual plays, per-model accuracy, chat Q&A, error-to-fix pairs, game outcomes, news sentiment, model-meeting insights, market text, and lifecycle evidence. Each table answers one recall question well.

Vector similarity indexingWhy this index

High-dimension embeddings exceed the standard HNSW limit, so vectors are indexed as half-precision with cosine ops. Approximate search trades recall, latency, and storage; validate the choice on representative queries.

Marker LedgerHow history survives

Every reset, graft, sweep and bundle writes one typed, additive row. Vectors answer "what is similar"; the marker ledger answers "what happened to this seat, and when." Nothing is ever mutated, so a reset is a time floor rather than a deletion.

Redis LayerWhy chosen

In-memory cache, cross-restart alert dedup, and signal pub/sub. Notifications survive restarts without double-firing, and hot data never touches the database.

Backup & ReplicationWhy chosen

Git history alone does not back up a database, secrets or runtime state, so backup and restore are documented separately and audited by a read-only readiness check. Recovery is a written runbook, not an assumption.

Retrieval DisciplineWhy chosen

Recall is filtered by distance thresholds and hit counters, cached answers are reused when close enough, and repeated model discussions are deduplicated before they waste tokens.

Design Rationale · 04

The engineering fleet: many concurrent AI sessions coordinated through a work ledger

Engineering changes benefit from isolated workspaces, explicit ownership, independent review, and reproducible checks.

The work logRead, append, close

Engineering changes benefit from isolated workspaces, explicit ownership, independent review, and reproducible checks.

Worktrees and pull requestsNo shared checkout

Engineering changes benefit from isolated workspaces, explicit ownership, independent review, and reproducible checks.

Agent wavesSix to ten in parallel

Engineering changes benefit from isolated workspaces, explicit ownership, independent review, and reproducible checks.

Incidents become guard testsFixed, then locked

Each significant incident is written down with the exact test that now enforces the fix, so a class of bug cannot quietly return. Automated remediation stays recommendation-only: an agent proposes, a human decides.

Design Rationale · 05

Reviewable skills and integration boundaries

This section illustrates component responsibilities and evidence boundaries. Private inventories, thresholds, and activation settings are intentionally omitted.

Deployment PackVercel

Preview snapshots, edge-config flags, dashboard deploys, and a release guard so the public surface ships through checks, never by hand.

Model Ops PackHugging Face

Shadow backtests, dataset curation, space evaluation runs, and a scheduled model scout feeding candidates to the bench.

Sports Data PackRotowire

DFS projections, injury pulse, lineup watch, and reanalysis triggers when ground truth changes under an open position.

Guardrail PackCashflow Guardrails

Drawdown circuit breaker, bankroll rebalancing, settlement idleness monitors, and watchdog operations for the money path.

Edge Intelligence PackDeep Picks

Early-edge alerts, confidence calibration, edge attribution, and a postgame loop that turns every result into analysis.

Readiness PackOps Readiness

Skill-coverage audits, plugin readiness orchestration, page release readiness, and a scheduled go/no-go verdict.

Knowledge PacksNotion · Dovetail

Roster sync, research databases, and project hubs in Notion; research memory, loss-cluster analysis, and postmortem mining in Dovetail.

Delivery PacksLinear · Resend · Cloudinary

Issue sync and SLO burn tracking, email briefs with delivery guards, and a chart CDN pipeline with visual audits.

Cloud Bridge PackGoogle Cloud

Warehouse queries, container deploys, storage artifacts, pub/sub bridging, managed inference, and secret rotation recipes.

Design Rationale · 06

Risk, safety & release discipline: engineered like it can fail

An autonomous system touching real markets gets the full defensive stack: position sizing with hard caps, multi-level kill switches, strict separation between simulated and real money, injection-scrubbed prompts, masked errors, authenticated surfaces, and a public boundary designed so leaking is structurally difficult.

Sizing DisciplineControl

Fractional Kelly sizing with a hard per-position cap, confidence-tier stake ladders, and minimum edge and expected-value floors before any stake is placed.

Hard BrakesControl

Automatic staking halts at the model, league, and model-league level when resolved win rates or P&L fall below floors. Plus a pause flag and a phone-side switch that stops the whole stack.

Paper-Real SeparationControl

Models trade virtual per-model bankrolls. Real accounts are tracked read-only, live actions require operator approval, and the browser execution bridge ships hard-capped and disabled by default.

Prompt SecurityControl

The design calls for constrained inputs, protected secrets, and reviewed publication boundaries. These mechanisms reduce risk; they do not establish complete prevention.

Secret HygieneControl

The design calls for constrained inputs, protected secrets, and reviewed publication boundaries. These mechanisms reduce risk; they do not establish complete prevention.

Authenticated SurfacesControl

Dashboards bind to localhost by default; remote access requires a token compared in constant time; write endpoints are gated and rate-limited.

Public Field WhitelistControl

The design calls for constrained inputs, protected secrets, and reviewed publication boundaries. These mechanisms reduce risk; they do not establish complete prevention.

Recovery DisciplineControl

Backup-readiness audits, fresh-clone recovery runbooks, machine-migration parity reports, and an append-only sync manifest make the whole system rebuildable from documentation.

Design Rationale · 07

One market, end to end

An illustrative path from a research question to a reviewed outcome.

Ingest

Markets stream in from the prediction venue through a whitelisted, deduplicated API client and land in Postgres with full provenance. Each of the four verticals has its own door, and every one of them fails closed on stale data.

Enrich

Sharp sportsbook lines, injury and lineup intelligence, live scores, websocket prices, and cited web research attach context to each market.

Recall

The vector brain surfaces similar past situations and their outcomes; the marker ledger contributes what has already happened to this seat, including any reset or graft.

Vote

The full panel runs the sport-tiered criteria rubric in parallel and each seat commits to a pick. A seat that does not answer is marked synthetic and excluded from every money query.

Grade

The golden-play contract decides whether this is a badge-worthy play: enough distinct voters, enough agreement, a licensed sample, and a published rate only ever beside the price-implied baseline.

Gate

Risk controls apply: confidence, edge, and expected-value floors, fractional Kelly sizing, stake ladders, hard brakes, and operator approval for anything real.

Resolve

Monitors track every open position, resolve outcomes API-first with model fallback, and stamp each prediction with its result.

Learn

Calibration updates, rankings move, embeddings absorb the outcome, and the substrate census records which learning stage actually ran. Where the evidence clears its floor, a losing seat is reset and re-taught the leading seat's method.

PolyMind · Public architecture showcase · Synthetic metrics only · No credentials, keys, or trading records