Research discipline is the moat, not a single winning number.
Most AI-trading products answer 'why should I trust your numbers?' with one pretty number. We answer by showing the verification process itself — and letting anyone check it again.
From signal to recommendation — the decision pipeline
Solid line = actually running in prod, verifiable. Dashed line = roadmap, not built yet — marked clearly so the two are never confused.
Not signal-push — a fixed cron schedule that automatically re-analyzes every asset hourly.
scheduler.py:533-539 — CronTrigger(minute=5), tasks.py:724 — run_brain_hourly_analysis()
Macro context (FRED, DXY, leading indicators) is injected into every analysis unconditionally. Net positioning (COT) currently applies only to metals, energy, and agriculture — VN stocks don't have this branch yet.
reasoning_engine.py:415-443 (COT), :449-493 (macro)
Wire COT/institutional positioning into the VN-stocks branch — currently only metals/energy/agriculture have it, a real gap, not something already done.
Since 2026-08-04, hourly gold and silver analysis has run through an automated debate round (one side argues bullish, one bearish, one judges) before a recommendation is recorded. Since 2026-08-08, Brain AI signals (Entry/SL/TP sent via VIP Telegram) also run through that same debate round BEFORE sending — before that date they did not, and this page used to say exactly that. If the judge leans against the signal's direction, we still send it and attach a 'CONTRARIAN' flag at the top of the message so you can decide for yourself. If the debate round CANNOT run (all model providers are unresponsive), the signal is still sent but carries a 'NOT DEBATED' label — on Telegram and on the card at /signals; we don't silently drop the signal, and we don't hide that the step didn't run. ON LATENCY, plainly: a signal only leaves the system AFTER the panel finishes, so it reaches you later than the moment it was detected — normally a few seconds, but when the free-tier models run out and the system has to fall back to paid models, it has measured up to ~46 seconds (measured in production 2026-08-16: 10.7s + 9.2s + 25.9s across the three seats). We choose to wait it out rather than cut the debate round short to send sooner. VN stocks and other assets don't have this step yet.
tasks.py::_run_ssm_debate_gate (ADI-1649, ADI-1754) + reasoning_engine.py:341-344 — DebateEngine
Reuse the exact 4-voice mechanism (Groq/Gemini/DeepSeek/OpenAI) currently gating the weekly newsletter — wire it into the hourly pipeline so every recommendation goes through debate before publishing, not just the weekly newsletter and not just gold/silver.
panel_debate.py — the mechanism exists; since 2026-08-16 (ADR-0014 slice 2b) it also gates brain_confluence_webhook.py (multi-asset confluence signals, including VN stocks) — no longer just metals_week_ahead_runner.py. The hourly Entry/SL/TP branch in step 3 still uses its own separate DebateEngine, NOT yet routed through panel_debate.
Recommendations are stored, timestamped, and entered into the public scoring loop.
prediction_service.py:50 — create_prediction()
Steps 2b and 3b are roadmap, not built — listed because they're the real direction, not because they're done. Full code trace: docs/product/business/RIGOR-AS-PROOF-ROADMAP.md §8 (baselify-brain). Re-checked against live code on 2026-08-27.
How confidence is measured
Every Brain recommendation carries a confidence label. A label is only worth trusting if it matches real outcomes, so Brain re-counts its closed signals per label and publishes the count here.
Recent examples
A data asset nobody else can copy
Not a price signal — a DECISION flow: whether you agree or disagree with a specific piece of reasoning, recorded BEFORE the outcome is known.
We record the reaction, not just the signal
Every recommendation on /signals has 3 choices: Agree / No / Ask more. Tapping 'Agree' or 'No' writes one decision row (the brain_signal_portfolio_decisions table) tied to that exact signal — at the moment you tap, the signal's outcome isn't known yet. A second flow — track 'in' / 'pass' (brain_signal_user_tracks) — records a reaction even when you hold no position to diff a weight against. Once a signal closes with a real outcome, both flows are reconciled against it.
Why a brokerage's order flow can't rebuild this
Order flow tells a brokerage what you traded. It doesn't tell them whether you believed the specific reasoning behind that trade — and there is no way to reconstruct, after the fact, a record of 'agree/disagree with THIS EXACT reason,' locked in BEFORE the outcome was known. That data only exists if the product deliberately asks the right question at the moment a recommendation is made, then writes it append-only — unchangeable once the outcome is known. Nobody can copy a history that has already passed.
Raw counts
Not enough sample yet to publish numbers — the threshold was set before measuring: at least 30 settled decisions, and each branch (agree / disagree / ask more / track-in / track-pass) at least 10. We don't show a small or stale number to avoid misleading you.
See the real interface
Illustrative example, fixed data — this page has no logged-in reader, so there's no real portfolio to diff a weight against. The live version runs on /signals for signed-in readers.
Portfolio proposal
Illustrative exampleWhat it means for your money
If it hits the stop level
≈ 2.4% of position
This is arithmetic from your position size to the stop level — NOT a profit forecast.
Closed loop: find the bug → fix it → re-verify on schedule
On 2026-07-28, a code-level audit (not just checking what's displayed) found 1 CRITICAL bug in institutional-positioning data — the COT column was reading the wrong CFTC report type, off by roughly 2x — plus a few HIGH-severity bugs.
Found
4 adversarial-audit sessions read the entire collector codebase, cross-checking live data against external CFTC/BlackRock/COMEX/Yahoo sources.
Fixed — 8/8
DINTG (5 tickets) + CHRD (3 tickets) — all Done, including a refactor into a shared adapter and a strict test suite.
Multi-model adversarial review — running for real, not a demo
4-voice debate gate
Every weekly market newsletter passes through Groq (bear case) / Gemini (missed factors) / DeepSeek (short vs. long horizon) / OpenAI (verdict) before publishing. A hard-fail gate, not an optional advisor.
War Room — Copilot
Designed to refuse to output a probability when there isn't enough sample size, instead of guessing just to show something. Running in prod.
Both have had real operational incidents (a model provider going down mid-run) that were handled transparently and logged — evidence this is a living system.
In June 2026, after the old approach failed the company's own self-imposed 'edge' test, leadership banned internally: no 'X% win rate', 'Sharpe > X', or 'beats the market' language on any public surface — even when it would help marketing. Very few retail fintechs adopt this policy before being forced to.
Active research
Not a product yet — stated upfront.
Measurement threshold locked 2026-08-02, no result yet. Current direction: converging evidence used to gradually size positions — not as a binary signal filter, a lesson learned after the old approach (a binary AND gate) was shut down internally by its own statistical evidence.
We deliberately don't list anything else here. Any feature without real data behind it doesn't appear in this document — no matter how appealing the name sounds.
Data source: HYPOTHESIS-REGISTRY.md (recounted per §1.1, verified twice, updated 2026-08-21) · RIGOR-AS-PROOF-ROADMAP.md §7.1/§7.2/§8 · Linear ADI-90.
baselify-brain is an intelligence + workflow layer, NOT a securities firm or broker. All performance figures are illustrative for design purposes and reflect our commitment to publishing results even when they fall short. Content is for reference only, not investment advice.