How a market idea earns the right to be taken seriously
Written by the Capstone Metals research desk. Reviewed by Daniel Kenney, President. Last updated February 2026.
The division of labour: code counts, people decide
Every number in our research is produced by deterministic code that can be run twice and give the same answer: returns, rankings, trading costs, slippage, turnover, volatility, drawdown, Sharpe, Sortino, Calmar, exposure, statistical tests and calibration. Language models are used for what they are actually good at — reading source material, proposing a falsifiable hypothesis together with the null it must beat, writing a specification, and arguing against a conclusion. A model never calculates a result and never authorises anything. Promotion of any strategy is a human decision, recorded as a proposal with a name attached.
The sequence every hypothesis walks
Each stage is stored with its inputs, outputs, model, timestamps, cost and status. A stage cannot be skipped: the server refuses a transition whose predecessor has not passed, so a promising result cannot jump the queue to the front.
1. Observation
Write down what was noticed in the data or the world, before any explanation is offered.
Written or reviewed by a person or agent
2. Data quality check
Check the data itself: gaps, duplicates, revisions, and whether there is enough history to say anything.
Computed by code
3. Hypothesis
State a falsifiable claim together with the null hypothesis it has to beat, so failure is possible.
Written or reviewed by a person or agent
4. Strategy specification
Turn the claim into a machine-readable specification: universe, features, execution model, costs, limits.
Written or reviewed by a person or agent
5. Code and unit checks
Implement it in deterministic code and unit-test the arithmetic against known values.
Written or reviewed by a person or agent
6. In-sample backtest
Run it over the testing history with commission and spread charged on every rebalance.
Computed by code
7. Walk-forward
Fit on one window, trade the next, repeatedly, so no parameter benefits from knowing the future.
Computed by code
8. Locked holdout
Read the reserved period once. It was locked before testing started and is never used for tuning.
Computed by code
9. Robustness
Perturb parameters, cut the data into subsamples and regimes, and multiply the trading costs.
Computed by code
10. Red team
Argue against the result on purpose: look-ahead, leakage, survivorship, too many trials, capacity.
Written or reviewed by a person or agent
11. Risk review
Check exposure, concentration and loss limits against the stated risk policy.
Written or reviewed by a person or agent
12. Research report
Write up what was found, what was not, and what remains unknown — including the failures.
Written or reviewed by a person or agent
13. Promotion proposal
Ask a named person to decide. Nothing is promoted automatically, ever.
Written or reviewed by a person or agent
The four tests that kill most ideas
- Walk-forward, not hindsight
- Parameters are fitted on one window and traded on the next, repeatedly, moving forward through history. A rule that only works when you already know how the period ended fails here.
- A holdout that is genuinely locked
- A period of history is reserved when a strategy version is created, before any testing begins. It cannot be changed afterwards through the admin screens or the API, and it is read once. If a conclusion falls apart there, it falls apart.
- Costs charged at a realistic level
- Commission and half-spread are charged on every rebalance, and the whole test is re-run with costs multiplied to see how much of the result was ever real. Findings that only exist at zero cost are recorded as failures, and kept as failures.
- Counting how many times we tried
- The number of variants tested is stored alongside the result, and used to compute a deflated Sharpe ratio and an estimate of the probability of backtest overfitting. Trying two hundred ideas and reporting the best one is not research; it is selection.
Adversarial review, on purpose
Before a research report is written, a red-team pass looks specifically for the ways we might have fooled ourselves: information used before it was knowable, labels leaking into features, a universe that quietly excludes what failed, too many trials, parameters that only work at one setting, fills that would not have happened, position sizes larger than the market could absorb, and results that depend on a single market regime. Any one of these can block a strategy from progressing, and the block is recorded rather than argued away.
What we openly do not have
- No paid market-data vendor, order-book feed or option-chain data. Our history is what we have recorded ourselves, and it is still short.
- No live brokerage connection. Nothing in this platform can place a real order, and that is a build-level fact, not a setting.
- No claim that a tested strategy will work in future. Testing narrows the range of plausible explanations; it never establishes one.
Evidence and sources
- 1. Bailey & López de Prado, The Deflated Sharpe Ratio (2014) — Why a Sharpe ratio must be discounted for the number of variants tried before it means anything.
- 2. Bailey, Borwein, López de Prado & Zhu, Pseudo-mathematics and financial charlatanism (2014) — The probability of backtest overfitting, and the CSCV procedure we use to estimate it.
- 3. Harvey, Liu & Zhu, …and the Cross-Section of Expected Returns (2016) — The multiple-testing problem in published factor research; the reason we count trials openly.
- 4. Novy-Marx & Velikov, A Taxonomy of Anomalies and Their Trading Costs (2016) — How much of a documented edge disappears once realistic trading costs are charged.
Where this fits in the wider picture
Market Intelligence is the evidence layer. The wealth-protection side of Capstone is where those findings meet an actual plan — metals, retirement accounts and stewardship of what you already hold.
- The education library — money, debt and purchasing power, from first principles.
- Gold's place in the world economy — what the historical record does and does not show.
- Wealth protection in one place — how metals, advice and entities fit together.
Research and education only. Nothing on this page is investment advice, a recommendation to buy or sell any security or metal, or a forecast. No outcome is promised or implied. Simulated and historical results do not indicate future results, and any strategy discussed here may lose money. Speak with us about your own circumstances before acting: (800) 200-9553.
