M.

← Experiment 001: the live desk

The Lab · Research Proposal · Paper in progress

Agentic Trading Under Pre-Registration: Momentum, Sentiment and Volatility Targeting at Retail Scale

MTUTHUKO MNGOMEZULU · ISINKWAMTASE RESEARCH · KWATHEMA, SOUTH AFRICA · 2026 · DRAFT v0.3

Abstract. Most retail algorithmic traders lose money, and most published trading results are survivorship-biased marketing. This study asks whether an autonomous AI agent, governed by pre-registered promotion criteria, can manage a deliberately small real crypto account without falling into the failure modes that destroy retail accounts: overtrading, tight-stop bleed, euphoria entries and unsized volatility. We describe an agentic audition protocol in which candidate strategies are generated and evaluated by fleets of AI research agents, must survive fee-inclusive backtests across nine years and ten assets, walk-forward windows spanning five market regimes, and bootstrap Monte Carlo analysis — before touching real capital, which itself remains in a withdrawal-disabled exchange sub-account. Five audition rounds (200+ backtests, 21 strategy variants) are complete — including a forensic audit of a real high-frequency retail bot whose live month paid 14.6× more in fees than it earned — and no variant has yet earned promotion, which we argue is the protocol succeeding, not failing. Preliminary findings include a large positive effect of volatility-targeted position sizing (identical signals: −62% → +229%), a sentiment gate that neutralised the 2021 mania cold-start, the persistent failure of short-selling and tight-stop styles, and the result that trade frequency, not signal quality, is the dominant predictor of retail account ruin.

1 · Background and motivation

Three forces motivate this study. First, the documented failure rate of retail algorithmic and leveraged trading — regulators worldwide report the large majority of retail derivative accounts lose money. Second, the arrival of capable AI agents makes it cheap to generate plausible-looking strategies at scale, which makes disciplined rejection machinery more valuable, not less: the bottleneck is no longer ideas but honest evaluation. Third, almost all public evidence in this space is promotional; experiments that publish their losses are rare. This project runs the entire lifecycle — research, simulation, live trading — in public, from a township-founded lab, on an account small enough that its loss is survivable and its lessons are general.

2 · Research questions

  1. RQ1. Can any rules-based strategy, net of real fees, beat buy-and-hold across full crypto market cycles at retail scale?
  2. RQ2. How much of strategy performance is attributable to position sizing rather than entry/exit signals?
  3. RQ3. Does aggregated crowd sentiment (Crypto Fear & Greed Index) carry usable signal beyond price-derived features?
  4. RQ4. Do pre-registered promotion gates materially change which strategies reach live capital, compared with common practice (deploying the best backtest)?
  5. RQ5. Can a high-frequency technical-ensemble bot — intraday charts, multi-indicator scoring, Bayesian trade gating — achieve positive expectancy net of fees at retail scale, and does its live behaviour match a faithful backtest of its rules?
  6. RQ6. Does a coin's unit price or liquidity tier ("cheap, high-volume altcoins") offer a retail edge that the majors do not, under the same rules and the same gates?

3 · Hypotheses

4 · Methodology

Adversarial review of the harness itself. From Round 6 the backtest harness is treated as a hypothesis too. Each round's code is attacked by three independent reviewers with distinct lenses — lookahead, execution realism, data and universe integrity — and the best out-of-sample result is re-implemented from scratch without the backtesting library and compared trade by trade. Round 6's reviewers found that the benchmark had been credited with a final trading day the strategies were denied (worth eight points out of sample) and that the two engines ran at different exposures; both were corrected before the verdict was read, and the verdict did not change. The reproduction matched the library to four decimal places on every trade.

5 · Data

6 · Preliminary findings (Rounds 1–6)

#FindingEvidence
F1Fees and overtrading, not signal quality, destroy small accounts first.194-trade strategy burned ~27% of account in fees; 1,800+ trade tight-stop style lost 96%.
F2Short-selling failed in every configuration tested.Three rounds; e.g. two-sided regime flip −89% vs long-only −7% on identical filters.
F3Position sizing dominates signal choice.Identical entries/exits: −62% naive vs +229% vol-targeted, drawdown 97% → 51%.
F4Crowd sentiment helps exactly where theory predicts.Extreme-greed gate turned the 2021 cold-start window from −83% to +69%; no full-path effect.
F5Nothing tested beats holding across full cycles; edges are not yet statistically significant.Market +1,168% over 9y; best variant +229%, p = 0.47; promotion denied by pre-registered gate.
F6A live retail bot's fee bill can exceed its profit by an order of magnitude while looking "profitable" day to day.Audited production bot: 708 real trades in one month, net +$0.80, fees $11.65 (14.6× net); long book +$5.00, short book −$4.20.
F7Trade frequency, not signal quality, is the dominant predictor of ruin; selectivity only slows the bleed.Faithful port of the bot's rules: −96% to −100% in and out of sample (6,500–20,000 trades); at maximum gating still −57% on 12 trades/day. Gross edge ≈ 0 at every gate level, both regimes.
F8Unit price is not an edge. "Cheap, liquid" altcoins are a volatility tier, not an opportunity tier.18 sub-$1 coins, 11 pre-registered variants, 7 walk-forward windows: the two-coin incumbent beat 10 of 11 out of sample; the regime rule that made +119% on BTC/ETH in-sample made −55% on the alt basket.
F9Altcoin-basket returns cluster in one or two half-years; a rule positive in ≤ 4 of 7 windows is a start-date bet, not an edge.Best breakout variant: +43% in-sample from two windows (+32%, +81%) with losses in the other five; the basket itself went +100% → −47% → −48% across three consecutive half-years.

7 · Planned work

8 · Ethics and disclosure

The study trades only the author's own, deliberately small capital in a withdrawal-disabled sub-account. It is not investment advice, offers no managed product, and solicits no funds. Published results include losses. Nothing in this study constitutes financial services under South African law.

DRAFT v0.3 · 2026-09-22 · Live experiment and daily decision journal: mngomezulu.africa/lab/agent-trading