Changelog

Every major change, improvement, and fix — documented transparently. ProphetAI is built in the open; this page tracks how the system evolves.

v4.3.02026-08-21Latest

Opus 5 Performance Fix — Confidence Collapse & Over-Betting

Deep performance analysis of Opus 5's first week revealed structural issues: only 1 HIGH confidence bet in 7 days (vs 86% under Opus 4.6), 31% of bets placed at LOW confidence losing money, and £28.50 bankroll drawdown. This release addresses the root causes — restoring quality-over-volume discipline, learned wisdom from 4 months of production data, and proper Kelly consultation.

  • Fix
    LOW confidence tracked bets banned. Selections below 58% probability are now tips only — the tracked ledger exists to grow the bankroll, not to speculate. This single change would have prevented £13.81 in losses from 14 losing LOW-confidence bets.
  • Fix
    Counter-argument guidance removed. The previous instruction telling the AI that 'counter-arguments adjust by 3-5pp' was causing systematic confidence collapse — Opus 5 was downgrading every bet from HIGH to MEDIUM after finding any objection. The AI is now encouraged to call HIGH confidence when evidence clearly supports it.
  • Improvement
    Volume discipline enforced: maximum 2 tracked bets per run (3rd requires HIGH confidence). Opus 5 was averaging 5.6 bets/day vs Opus 4.6's 3.4 — more activity but dramatically lower quality.
  • Feature
    Learned Wisdom section added to all prompts — 4 months of hard-won insights from ~500 tracked bets and ~2900 tips. Includes MLB pitcher-edge thresholds, DNB/spread outperformance data, sport-specific anti-patterns, and process lessons. Knowledge transfer without reintroducing the problematic self-learning loop.
  • Improvement
    Kelly consultation now mandatory before stake sizing. The AI can still deviate but must articulate why in its rationale — preventing the silent overexposure that was burning bankroll on thin edges.
  • Feature
    Rolling outcome feedback: the AI now reviews its last 10 settled bets (with W/L results, confidence, and P&L) at the start of each run. Provides a non-permanent learning signal — enough context to self-correct without accumulating the prohibitions that paralysed previous strategy files.
  • Improvement
    Mission reframed from 'the product must feel alive' to 'three good bets that win beats ten with a 70% loss rate'. Tips are now explicitly positioned as the content engine; tracked bets as the precision bankroll instrument.
v4.2.02026-08-05

Confidence Calibration & Tool Pipeline Overhaul

Major overhaul of the AI's confidence calibration system and tool pipeline. Fixes the systematic confidence anchoring at 60%, restructures when and how the AI reads its own performance data, and replaces fear-inducing tool outputs with neutral factual language.

  • Fix
    Fixed confidence anchoring at 60%. The per-tier deflation system created a cliff edge where MEDIUM confidence always produced minimum stakes — the AI learned to park at 60% because it was the 'safe' zone. Replaced with unified deflation that applies equally across all confidence levels, creating a smooth staking gradient.
  • Architecture
    Tool pipeline restructured. Calibration and EV analysis tools moved from pre-research (step 1) to pre-staking (step 5). The AI now researches events on their merits first, then consults its track record when sizing stakes — not the other way around.
  • Fix
    Restored Discord rejection posting ('The Ones That Got Away'). Rejections hadn't been posting since July 19th due to Codex JSONL output format not being parsed correctly. The orchestrator now handles both raw text and JSONL transcript formats.
  • Improvement
    ROI target lowered from an unrealistic 30% to 5%. The old target triggered a permanent CONSERVATIVE stance on every single run — the AI was told it was failing before it even started researching.
  • Improvement
    Calibration and deflation now use rolling windows (20–50 recent bets) instead of all-time data. Early mistakes under the broken strategy loop no longer permanently poison the AI's staking decisions.
  • Improvement
    Tool outputs stripped of catastrophising language. 'Edge detection may be broken' replaced with factual variance context. Sport-avoidance rules removed — bad luck no longer permanently brands a sport as toxic.
  • Fix
    Prompt trimmed by ~140 lines per model. Removed hardcoded statistics, prescriptive rules, and contradictory advice. The AI is now clearly positioned as the decision maker with advisory tools, not a rule follower.
  • Feature
    New Wilson confidence intervals in calibration output — the AI can now distinguish statistically significant over/under-confidence from noise. New EV realisation CLI tool for detailed edge analysis.
v4.1.02026-08-03

Strategy Loop Removal & Confidence Fix

Removes the self-learning strategy loop that was causing agent paralysis, fixes confidence double-deflation, and replaces the AI Strategy section with live per-run decision insights.

  • Architecture
    Strategy self-learning loop removed. The AI no longer reads or writes strategy files between runs — eliminating the accumulated prohibitions that were causing decision paralysis across model generations.
  • Fix
    Confidence double-deflation fixed. The AI now reports its genuine probability estimate for each event. The Kelly/calibration system handles stake sizing independently, preventing the compounding deflation that was producing minimum-stake coin-flip bets.
  • Feature
    New 'Recent Insights' section replaces AI Strategy on the homepage. Each run now produces a short, punchy decision diary entry — what the market looked like, what caught the AI's eye, and why it backed or passed on events.
  • Improvement
    Run insights are write-only snapshots stored in the database — the AI cannot read them back and spiral into self-doubt. Previous strategy files remain on disk but are dormant.
v4.0.02026-07-20

Multi-Model Architecture & GPT-5.6 Sol

ProphetAI now supports multiple AI model generations with full attribution, independent strategy evolution, and generation-aware performance tracking. GPT-5.6 Sol takes over from Claude Opus 4.6 as the live model.

  • Feature
    Multi-model architecture: every bet, tip, and acca is now attributed to the specific AI model that generated it. Each model generation gets its own bankroll, strategy, and performance metrics — enabling clean comparisons and seamless future model upgrades.
  • Feature
    GPT-5.6 Sol introduced as the new live model. Sol starts with a fresh £100 bankroll and develops its own strategy from scratch, independently of Opus's 88-day track record.
  • Feature
    Website redesigned for generation-aware reporting: stats page shows per-model performance with tabbed breakdowns, a model comparison table with separate bet and tip statistics, and generation-filtered bankroll charts. The homepage shows only the live model's P&L and bankroll.
  • Feature
    Model transition banner: when a new model takes over, a prominent banner appears explaining the change and linking to historical performance. Open positions from previous models are clearly marked with attribution pills.
  • Improvement
    Each model now evolves its own strategy independently — no cross-contamination between generations. New models start clean without inherited biases from previous iterations.
  • Fix
    "STILL IN PLAY" Discord replies can now reopen incorrectly settled bets and tips, fixing a long-standing issue where premature settlements couldn't be reversed.
  • Improvement
    Confidence calibration and risk management tools are now model-aware, so new models receive relevant advice based on their own performance history.
  • Improvement
    Market search window standardised to a consistent rolling +48h window, ensuring all daily runs get equal coverage regardless of time of day.
v3.5.0July 2026

Market Expansion & Strategy Intelligence

Major expansion of available markets across all sports, deeper exchange coverage for niche sports, and a fundamental rethink of how the AI learns from its own reasoning.

  • Feature
    All non-soccer sports now receive spreads and totals data alongside moneyline — previously limited to h2h only. The AI can now identify value in handicap and over/under markets across basketball, baseball, ice hockey, American football, tennis, cricket, and more.
  • Feature
    Exchange market coverage expanded from match odds only to sport-specific market types across 18 niche sports — including handicaps, totals, method of victory, round betting, correct score, and placement markets for sports like darts, snooker, boxing, MMA, cricket, rugby, and golf.
  • Architecture
    Strategy files converted from stat-heavy to purely qualitative guidance. The AI now queries its own database for live performance data every run rather than relying on stale cached numbers — ensuring decisions always reflect current reality.
  • Feature
    New reasoning retrospective system: the AI reviews its own research process and rationale for recently settled bets, correlating how it thought with what actually happened. This builds methodology insights over time — not just 'what won' but 'what thinking process leads to good decisions'.
  • Settlement
    Tennis game markets (spreads and totals based on games, not sets) now correctly use specialised settlement logic, since standard score APIs return set counts that don't work for game-based betting lines.
  • Fix
    American football is no longer incorrectly treated as soccer — NFL overtime games previously risked being settled using 'regulation time only' rules intended for soccer's extra time.
  • Settlement
    Full integration test coverage added for NBA, MLB, NHL, and NFL spread and totals settlement — 19 new tests ensuring every major US sport settles correctly across all market types.
  • Improvement
    Website performance display reframed: 'Last 7 Days' and 'Previous 7 Days' rolling windows replace 'This Week / Last Week' for precise, non-misleading labelling. Status pills (Hot Streak, Accelerating, etc.) added to homepage cards.
v3.4.0June 2026

Multi-Currency & Performance Transparency

Full multi-currency localisation across Discord and website, plus a rethink of how performance data is presented to tell the story of recent improvement.

  • Feature
    Multi-currency support: all financial values can now be viewed in ~30 currencies. Discord embeds include a 'View in my currency' button, and the website auto-detects your locale with a manual override panel.
  • Feature
    Performance display overhauled — recent 7-day and 14-day windows prioritised over all-time stats on both the homepage and stats page, showing the bot's trajectory rather than just lifetime averages.
  • Feature
    Stats page redesigned: separate tip performance section with trend graphs, sport-by-sport breakdown with timeframe clarity, and promotional banner showing hypothetical returns when the story is positive.
  • Feature
    Tip-ACCA cross-references: individual tips are now linked to their parent accumulator (and vice versa), enabling the website to show which tips form part of accas.
  • Improvement
    Bet history now sorted by settlement time (not placement time) across all views — form indicators, history page, and API — giving a consistent and accurate picture of recent results.
  • Fix
    Acca odds excluded from ledger average odds calculations — only actual bet odds are used, preventing misleading inflation of displayed average odds.
v3.3.0May 2026

Turnaround Engine

Data-driven qualification scoring, tiered governance that scales with performance, and tools to replay and learn from historical decisions.

  • Feature
    Qualification Scorecard: every tracked bet candidate is scored 0–100 across six factors (closing-line value, odds movement, segment record, research depth, liquidity, and recency). Below-threshold picks are blocked or flagged.
  • Architecture
    4-tier governance system (Survival / Recovery / Growth / Normal) with tier-specific Kelly fractions, daily caps, stake limits, and qualification thresholds. The system automatically loosens restrictions as performance improves.
  • Feature
    Replay harness for backtesting: chronological replay of historical events through the current qualification and governance system, producing Brier scores, calibration data, and simulated P&L.
  • Feature
    Calibration visualisation: predicted vs actual win rates, broken down by confidence tier and sport segment. Available as both interactive JSON and exported PNG plots.
  • Improvement
    Run cadence reduced from 7x to 5x daily — better quality decisions with less data cost. Agent philosophy shifted to 'tools advise, hard rails protect, agent decides'.
  • Fix
    Calibration deflation capped at 15pp with a rolling 50-bet window — prevents death-spiral where a few bad bets cause the system to over-correct and refuse to place any picks.
v3.2.0May 2026

Bankroll Rescue

Emergency governance layer to arrest bankroll decline caused by AI overconfidence and excessive volume. Hard mechanical gates that the AI cannot override.

  • Architecture
    New bankroll governance layer: hard blocks on MEDIUM confidence tracked bets, daily bet caps, exposure limits (50% max), and minimum odds thresholds (1.50). The AI selects what to bet — governance decides whether it may touch bankroll.
  • Fix
    Recommended ACCAs blocked by default after 1W-9L record and -25% ROI. Re-enabled later with proper advisory guardrails and conservative sizing (1-2% of bankroll).
  • Improvement
    Prompt philosophy rewritten: removed volume pressure and daily quotas, replaced with quality-over-quantity messaging. Tips now processed after tracked picks to prevent duplicate detection from blocking stronger picks.
  • Fix
    Tips were blocking tracked bets via the duplicate detection system — tips are now allowed to be promoted to tracked picks without triggering dedup.
  • Improvement
    Event window expanded from 48 hours to 72 hours — gives the AI more time to find value in upcoming fixtures rather than rushing decisions.
v3.1.0May 2026

Credit Leak Fix & Self-Learning

Critical fix for a data credit leak that was burning through the monthly quota, plus the foundations of the bot's self-learning system.

  • Fix
    Fixed critical data provider credit leak in settlement that was consuming ~14,400 unnecessary credits per day. Accurate credit tracking now uses response headers instead of estimates.
  • Feature
    Quota monitoring: Discord warnings at 50%, 60%, 70%, 80%, 90%, and 100% of the monthly budget. Prevents surprise overages and lets us adjust run frequency proactively.
  • Feature
    Self-learning foundations: performance reports, calibration analysis, strategy rule scoring, and Kelly-validated staking with confidence deflation. The bot can now assess its own accuracy by confidence tier.
  • Improvement
    Settlement efficiency: event-time gating prevents premature lookups, 60-minute TTL cache eliminates redundant queries, and enrichment regions narrowed to UK-only (halving per-event data cost).
  • Feature
    AI-controlled Kelly fraction: the bot can adjust its own staking aggressiveness within guardrails, with Discord notifications when it changes its approach.
v3.0.0May 2026

Multi-Agent Pipeline

Complete architecture overhaul — from single-shot agent to multi-stage pipeline with deterministic data gathering and structured decision-making.

  • Architecture
    New 2-stage pipeline: deterministic data gathering (Python, no AI cost) followed by a unified LLM agent session for shortlisting, research, and decisions. Dramatically reduces cost per run while improving data quality.
  • Feature
    Upgraded to a comprehensive primary odds provider — covering 30+ sports with 40+ bookmaker odds for every event. Previously limited to a single-market view.
  • Feature
    Discord embeds now show Best Odds (Top 5) comparison for all bets and tips, letting users find the best price across bookmakers rather than relying on a single source.
  • Feature
    Research cache with TTL and enrichment deduplication — prevents redundant data calls across runs and ensures the AI always works with fresh information.
  • Improvement
    Pipeline health alerting: consecutive failures trigger warnings, and a cost tracking ledger monitors data credits per run to prevent budget overruns.
  • Settlement
    Tiered settlement upgraded with additional score validation: cross-referencing multiple sources to prevent settling on incorrect or partial scores. Sport-specific time buffers for cricket (4h), baseball (4h), and basketball (3h).
v2.0.0April 2026

Exchange Integration & Settlement Hardening

Exchange integration for racing and niche sports, accumulator support, and a complete overhaul of the settlement infrastructure.

  • Feature
    Exchange API integration: live exchange odds, racecards, and authoritative settlement for horse racing, greyhound racing, and motor sport. Per-session API call caps and steward's enquiry awareness built in.
  • Feature
    Accumulator support: new #accumulators Discord channel with treble and 4-fold+ accas. Void/push legs reduce the acca using standard bookmaker rules. Full settlement pipeline for individual legs and parent acca status.
  • Settlement
    Multi-tier settlement pipeline with multiple independent validation sources, AI web research fallback, and human review via Discord. Settlement decoupled from agent runs into a 15-minute independent loop.
  • Fix
    Settlement defences: time guards prevent premature settlement, series metadata prevents cross-match confusion (e.g. MLB series games), and full audit trails for every outcome change.
  • Architecture
    Dev/production environment split: separate git worktrees, databases, Discord channels, and configs. Guarded promotion flow ensures tested code before production deployment.
  • Fix
    Score cross-validation catches wrong scores from upstream providers. When sources disagree on the outcome, the more authoritative source is used and discrepancies are logged for review.
  • Settlement
    Human-in-the-loop settlement: overdue items (>24h unfindable) post to a dedicated #discrepancies channel. Human replies (WON/LOST/VOID) are parsed and applied automatically, with override audit trails.
v1.0.0April 2026

Initial Launch

The beginning — a virtual bankroll AI sports analyst posting picks to Discord with a transparent website dashboard and the foundations of self-improving strategy.

  • Feature
    Core bot: AI-powered sports analysis agent running multiple times daily, researching live events, and posting tracked picks and tips to Discord channels with full rationale for every decision.
  • Feature
    Virtual bankroll system: configurable starting bankroll with Kelly-informed stake sizing, P&L tracking, and automatic daily stats aggregation. No real money wagered — ever.
  • Feature
    Discord integration: sport-specific channels (football, tennis, horse racing, etc.), rich embeds with odds, confidence level, and event details. Separate tracked picks (bankroll impact) from tips (signal only).
  • Feature
    Website dashboard: live performance display with bankroll chart, win/loss record, ROI tracking, full bet history table, and strategy view — all synced automatically from the bot's database after each run.
  • Feature
    Strategy evolution: the AI maintains and updates its own strategy document after each run based on outcomes, codifying patterns with sufficient sample sizes and deprecating rules that stop working.
  • Settlement
    Basic settlement system: automatic score lookups with time-guarded settlement windows. Bets that can't be auto-settled are flagged for manual review.
  • Architecture
    SQLite database with schema migrations, Discord message tracking, and deterministic stats computation. Blog/diary system where the AI writes about its own performance and market observations.

ProphetAI has been in continuous development since April 2026. This changelog covers major releases — minor patches and hotfixes are applied between versions.