Capstone Metals — gold and silver IRA dealer
Speak with a specialist: (800) 200-9553Monday – Friday, 7am – 4pm Pacific

Quantitative research

How a market idea earns the right to be taken seriously

Most published market findings do not survive contact with honest testing. That is not a cynical remark — it is the documented result of the multiple-testing literature. This page describes the machinery we built to make our own ideas fail early, cheaply and in writing, rather than expensively and in public.

Written by the Capstone Metals research desk. Reviewed by Daniel Kenney, President. Last updated February 2026.

The division of labour: code counts, people decide

Every number in our research is produced by deterministic code that can be run twice and give the same answer: returns, rankings, trading costs, slippage, turnover, volatility, drawdown, Sharpe, Sortino, Calmar, exposure, statistical tests and calibration. Language models are used for what they are actually good at — reading source material, proposing a falsifiable hypothesis together with the null it must beat, writing a specification, and arguing against a conclusion. A model never calculates a result and never authorises anything. Promotion of any strategy is a human decision, recorded as a proposal with a name attached.

The sequence every hypothesis walks

Each stage is stored with its inputs, outputs, model, timestamps, cost and status. A stage cannot be skipped: the server refuses a transition whose predecessor has not passed, so a promising result cannot jump the queue to the front.

  1. 1. Observation

    Write down what was noticed in the data or the world, before any explanation is offered.

    Written or reviewed by a person or agent

  2. 2. Data quality check

    Check the data itself: gaps, duplicates, revisions, and whether there is enough history to say anything.

    Computed by code

  3. 3. Hypothesis

    State a falsifiable claim together with the null hypothesis it has to beat, so failure is possible.

    Written or reviewed by a person or agent

  4. 4. Strategy specification

    Turn the claim into a machine-readable specification: universe, features, execution model, costs, limits.

    Written or reviewed by a person or agent

  5. 5. Code and unit checks

    Implement it in deterministic code and unit-test the arithmetic against known values.

    Written or reviewed by a person or agent

  6. 6. In-sample backtest

    Run it over the testing history with commission and spread charged on every rebalance.

    Computed by code

  7. 7. Walk-forward

    Fit on one window, trade the next, repeatedly, so no parameter benefits from knowing the future.

    Computed by code

  8. 8. Locked holdout

    Read the reserved period once. It was locked before testing started and is never used for tuning.

    Computed by code

  9. 9. Robustness

    Perturb parameters, cut the data into subsamples and regimes, and multiply the trading costs.

    Computed by code

  10. 10. Red team

    Argue against the result on purpose: look-ahead, leakage, survivorship, too many trials, capacity.

    Written or reviewed by a person or agent

  11. 11. Risk review

    Check exposure, concentration and loss limits against the stated risk policy.

    Written or reviewed by a person or agent

  12. 12. Research report

    Write up what was found, what was not, and what remains unknown — including the failures.

    Written or reviewed by a person or agent

  13. 13. Promotion proposal

    Ask a named person to decide. Nothing is promoted automatically, ever.

    Written or reviewed by a person or agent

The four tests that kill most ideas

Walk-forward, not hindsight
Parameters are fitted on one window and traded on the next, repeatedly, moving forward through history. A rule that only works when you already know how the period ended fails here.
A holdout that is genuinely locked
A period of history is reserved when a strategy version is created, before any testing begins. It cannot be changed afterwards through the admin screens or the API, and it is read once. If a conclusion falls apart there, it falls apart.
Costs charged at a realistic level
Commission and half-spread are charged on every rebalance, and the whole test is re-run with costs multiplied to see how much of the result was ever real. Findings that only exist at zero cost are recorded as failures, and kept as failures.
Counting how many times we tried
The number of variants tested is stored alongside the result, and used to compute a deflated Sharpe ratio and an estimate of the probability of backtest overfitting. Trying two hundred ideas and reporting the best one is not research; it is selection.

Adversarial review, on purpose

Before a research report is written, a red-team pass looks specifically for the ways we might have fooled ourselves: information used before it was knowable, labels leaking into features, a universe that quietly excludes what failed, too many trials, parameters that only work at one setting, fills that would not have happened, position sizes larger than the market could absorb, and results that depend on a single market regime. Any one of these can block a strategy from progressing, and the block is recorded rather than argued away.

What we openly do not have

  • No paid market-data vendor, order-book feed or option-chain data. Our history is what we have recorded ourselves, and it is still short.
  • No live brokerage connection. Nothing in this platform can place a real order, and that is a build-level fact, not a setting.
  • No claim that a tested strategy will work in future. Testing narrows the range of plausible explanations; it never establishes one.

Evidence and sources

  1. 1. Bailey & López de Prado, The Deflated Sharpe Ratio (2014) Why a Sharpe ratio must be discounted for the number of variants tried before it means anything.
  2. 2. Bailey, Borwein, López de Prado & Zhu, Pseudo-mathematics and financial charlatanism (2014) The probability of backtest overfitting, and the CSCV procedure we use to estimate it.
  3. 3. Harvey, Liu & Zhu, …and the Cross-Section of Expected Returns (2016) The multiple-testing problem in published factor research; the reason we count trials openly.
  4. 4. Novy-Marx & Velikov, A Taxonomy of Anomalies and Their Trading Costs (2016) How much of a documented edge disappears once realistic trading costs are charged.

Where this fits in the wider picture

Market Intelligence is the evidence layer. The wealth-protection side of Capstone is where those findings meet an actual plan — metals, retirement accounts and stewardship of what you already hold.

Research and education only. Nothing on this page is investment advice, a recommendation to buy or sell any security or metal, or a forecast. No outcome is promised or implied. Simulated and historical results do not indicate future results, and any strategy discussed here may lose money. Speak with us about your own circumstances before acting: (800) 200-9553.

Call (800) 200-9553 Text Email