How the model was built, tested, and rebuilt.
The research process behind the Ticker Nerd Rank and the Ticker Nerd 20 — the data, the anti-curve-fitting rules, the test battery, and the strategies we killed on the way. The short version is first. The detail after it is for anyone, reader or machine, who would rather check than trust.
Updated 3 Aug 2026Ten facts that define the method.
- Every factor traces to published research. Value, momentum, quality and the rest were documented by named academics decades ago; none of the signals are in-house inventions. The papers are listed at the foot of this page.
- Around fifty measurements, nine families, weights close to even. No weight was ever optimised — leaning hard on one factor is a bet on which will work next, and the weights stay near even instead.
- The data is point-in-time. More than twenty years of Compustat and FactSet history from 2005, showing the model only the preliminary figures an investor could have seen on the day rather than the restated ones published later.
- Three locked windows. Rules were tuned on 2005–2014, checked on 2015–2020, and judged once on 2021–2025 — a blind window that is spent the moment it is read.
- Earlier strategies failed that blind test and were killed. They had passed every in-sample check first. We did not re-tune them; we rebuilt simpler, with humbler targets — and we count the rebuilt model's result on those years as a post-development evaluation, not a blind pass.
- One theory-justified value per parameter. Holding count, review cadence and every threshold were set on economic reasoning before testing — never by running variants and keeping the winner.
- Hundreds of stress tests. Rolling start dates, seven market regimes, weight and universe perturbations, turnover checks — judged on the worst runs rather than the average.
- Costs are modelled per stock. Commission plus slippage scaled to each stock's actual liquidity, with fills at the next session's average price rather than the close the signal used.
- Rejections are logged. Factors that failed are recorded with the reason, and some famous ones are absent from the model because they failed here.
- The benchmark is the S&P 500 Equal Weight. The S&P 500 is always shown beside it, and the Nasdaq 100 is used to check the model is not a growth bet in disguise.
This page is also one plain markdown file. Copy it into your favourite AI and ask for a summary in your terms — or ask it to argue the other side.
The rules were set before the tests were run.
A backtest can be made to say anything. Test enough variants and one will look brilliant by chance — and then fail with real money. Every rule in this section exists to close that road.
One theory-justified value per parameter. Each choice — a factor's definition, a threshold, the holding count — gets one value, argued from economics before any test runs, and the verdict is in or out. We do not try ten variants and keep the best-looking.
Sweeps check for plateaus, not peaks. When a setting is varied around its chosen value, the question is fragility. A result that only holds at one precise setting is treated as luck, and the baseline never moves to the peak.
Economic sense outranks the backtest. A signal must have a working mechanism — the three-part test in the factor-selection section below — or it is presumed to be data-mined noise, however good it looks. That is the kind that fails out of sample.
Tests are judged on the bad runs. Headline return is a spotlight number. Gates sit on the fifth-percentile outcome across hundreds of windows, because the bad stretch is the one an investor actually has to live through.
The blind window is spent the moment it is read. Out-of-sample data you have peeked at is in-sample data. Once the blind years delivered a verdict, tuning against them was forbidden — a new idea needs new, unseen data.
Broad, liquid, and hard to game.
The research universe is roughly the 1,500 largest, most liquid US-listed companies, rebuilt by rule at every date across the whole tested history. The portfolio only ever chooses from names liquid enough to trade at the costs the simulation charges.
Ranks outside that set are informational. The tested history covers the ~1,500 investable names. The rest of the ~4,600 the site scores are ranked by the same arithmetic, but they sit outside the validated universe — smaller companies carry patchier data and behave differently, so a high rank there is a weaker piece of evidence.
Financial companies stay in. Banks and insurers break many standard measurements — a bank has no gross profit — and the easy fix is to exclude the whole sector. We measure them on bank-appropriate terms instead, because a universe rule that quietly deletes a fifth of the market is a hidden sector bet.
A bug taught us that rule. An early profitability screen ranked a missing field as the worst possible score and silently excluded every bank in the universe — and the backtest looked better for it. Finding that bug, and fixing it by measuring financials on their own terms, shaped how every later factor was built.
No domicile filter. A draft rule excluded companies incorporated abroad, which turned out to remove US-listed multinationals like Linde and Accenture. The rule was measuring paperwork, so it was cut.
Data problems are fixed where they occur. When a data quirk appears, the fix goes into how the affected measurement is computed. Cutting the universe to dodge bad data shrinks the opportunity set for every factor at once.
The eligibility rules are public. In the order they run:
- Common shares on a US exchange, primary listing only — no ADRs, no partnerships (MLPs), no over-the-counter listings, no preferreds or warrants.
- Standard sectors only: the sparse catch-all sector buckets are out, and so are REITs — trust accounting (mandatory payouts, funds-from-operations) breaks the earnings- and accrual-based measurements, and unlike banks there is no fair same-scale substitute.
- Share price above $5 — a penny-stock guard that almost never binds at this size.
- Median daily traded value above $10 million over the trailing quarter — the binding investability gate. A median rather than an average, so one spike of volume cannot carry an illiquid name in.
- Trailing twelve-month sales above zero — no pre-revenue shells or SPACs.
- No announced takeovers — once a deal is public, the price is pinned to its terms.
- At least a year of trading history — young IPOs distort momentum and volatility measurements.
- Then the 1,500 largest companies that survive all of the above. The order matters: eligibility first, size rank last — the same construction logic the index providers use.
A factor must explain why it should keep working.
The bar is a mechanism. A return premium must be compensation for one of exactly three things: a risk others will not bear, a persistent behavioural mistake, or a structural friction. A signal that fits none of the three is presumed to be noise, whatever its backtest says.
Measurements climb the income statement. The higher a number sits, the harder it is to manage: gross profit is harder to game than operating profit, and both are harder than adjusted earnings. Where possible, the model measures near the top — the reasoning behind Novy-Marx's gross-profitability work.
A missing number is not a zero. Whether a blank field means none, unreported, or does-not-apply changes what a measurement says about a company, so every measurement defines its missing values explicitly.
The steepest backtest curves get extra suspicion. Signals with the most dramatic in-sample results tend to be crowded, regime-specific, or fitted to a few lucky periods — exactly the kind that fades when real money arrives.
Some measurements compare a company against its own industry rather than the whole market, where the economics differ too much for a raw comparison — a software company and a steel mill hold assets for different reasons.
Construction is judged before the signal is. One factor looked dead until its raw count was adjusted for how many analysts follow each stock; it then became one of the stronger signals in its family. A measurement is only condemned after its construction is right.
A candidate must beat its own shuffled ghost. Before a new measurement counts as signal, its marginal contribution is compared against a permutation noise floor — the same addition with its per-stock values shuffled. An improvement the shuffle can match is noise.
Nine families, weighted close to even.
The list below reads live from the ranking system — the same tree that scores every stock daily.
Around fifty individual measurements roll up into nine families. The weights sit close to even on purpose, and none of them was set by optimisation: every factor has had bad years, and leaning hard on one is a guess about which works next.
One deliberate tilt is size. Smaller companies score higher, all else equal — fewer professionals study them, so they are more often mispriced. It is why the portfolio holds names most members have never heard of.
The factor mix itself tips the book aggressive. The weights sit near even, but which factors made the list reflects the portfolio's mandate. Growth is the plainest example: on its own it tested as a poor predictor, yet alongside the other families it added value — and that combination test, never its solo backtest, is why it is in.
The taxonomy is public; the implementations are the product. The exact formulas, comparison scopes and weights are the engine members pay for, and publishing them would hand out the recipe without adding evidence. The tests below, and the live record, are the checkable part.
Which way analysts are moving their numbers, and how fast.
What the shares cost against the profits, cash flow and assets standing behind them.
How the shares have performed over the past months, next to every other stock.
How much profit the business earns on the capital it uses, and whether it holds up year after year.
How much of the reported profit actually arrives as cash, and how fast the company is expanding its asset base.
Whether the share count is growing or shrinking, and how much the company leans on outside financing.
Which way sales and profits are heading, and whether the pace is picking up or fading.
How large the company is against everything else the model tracks.
How violently the share price moves in ordinary weeks.
What each factor means, in plain English — the full glossary, with the research behind each family. As at 1 Aug 2026.
The rejections are how you know the survivors were tested.
Every rejection is logged with its reason. A model that only shows what made it in is showing half the work. This section is the half that usually stays hidden.
Price-to-book — the most famous value measure — is absent. Across twenty years of our universe it carried no useful signal, and a version adjusted for intangible assets tested no better. The model's value family is built on earnings, cash flow, sales and payouts instead.
Market timing is absent. A trend rule that steps aside when prices fall below their long-term average tested badly: it sells after falls and re-enters after recoveries, missing the rebounds that make long-term returns. The model stays invested.
Analyst rating changes are absent; forecast changes are in. Upgrades and downgrades tested dead in two separate constructions. What analysts do to their numbers carries information; the rating labels arrive too late.
Insider buying is absent. In our universe and holding period it pointed the wrong way in testing — a finding that surprised us, and stood.
Short-term reversal and seasonality never entered. Reversal turns over far too fast for a portfolio reviewed every four weeks, and seasonality is missing from one of the two major replication studies — two independent literatures disagreeing was reason enough to skip it.
Whole architectures were rejected too. Separate single-style portfolios run side by side proved too correlated to add anything, and averaging their scores into one rank picked pleasant all-rounders that topped no list. The integrated hierarchy above replaced both.
Twenty years of history, hundreds of ways to fail.
The engine is an institutional-grade backtesting platform running Compustat and FactSet data — the same databases professional managers buy.
Every test is point-in-time. On any past date the model sees the preliminary figures available that day rather than the cleaned-up numbers restated later. A backtest without that discipline grades itself with tomorrow's answer sheet.
Three windows, locked in advance. 2005–2014 to build, 2015–2020 to check, 2021–2025 held blind for a single final verdict. Contamination is policed: one strategy's blind result is never used to tune a sibling.
Start-date luck is tested directly. Each configuration runs across hundreds of overlapping multi-year windows, offset a week at a time, and is gated on its fifth-percentile outcome — the run where the timing went against it. The windows overlap heavily, so they are never counted as hundreds of independent observations: this test measures sensitivity to start date, not statistical significance.
Seven named regimes. Every candidate is scored separately through the pre-2008 boom, the financial crisis, the recovery, the mid-cycle years, the COVID crash, the 2022 rate shock, and the years since. A rule that only works in one kind of market fails here.
Construction is perturbed around every chosen value. Holding count, review cadence, family weights nudged up and down, each family dropped in turn. The gates are one-sided — only collapses fail — and a strong result at a neighbouring setting never moves the baseline.
The universe is perturbed too. Fewer names and more, higher and lower liquidity floors, financials out — checking the result is a property of the method rather than of one particular list of stocks.
Tail dependence is measured. If deleting the five best trades guts the result, the strategy was five lucky trades. A result has to survive losing its own highlight reel.
Then the disguise checks. Returns are regressed on the five Fama-French factors plus momentum, using Kenneth French's public data, and the model is tested against a blend of low-cost factor ETFs. A strategy that a three-fund ETF portfolio can replicate has no reason to exist — and one whose edge is a bet on large growth stocks shows up immediately against the Nasdaq 100.
The research behind the current model left hundreds of saved test batteries, most spanning more than two hundred overlapping windows each. The count is not the point; the point is that no single flattering backtest ever decided anything.
The test that killed two strategies.
In 2026 the first two finished strategies met the blind window, and failed. Both had passed every in-sample tier — the rolling windows, the regimes, the perturbations. Judged on five years they had never seen, the edge was gone.
They were killed, and never re-tuned. Adjusting a strategy to fix its blind years turns the last clean data into more training data, and the next backtest into fiction. The discipline held; the strategies did not.
The post-mortem said why, and the rebuild went simpler. The failed versions had leaned on the signals that looked steepest in training — the crowded, regime-fitted kind. The rebuild stripped those out rather than adding cleverness: fewer moving parts, weights spread close to even, and a return target cut to something the evidence could carry.
A window read once is degraded for every model that comes after. Having watched two strategies fail on 2021–2025, we knew what failure there looked like — so although the rebuilt model had never touched those years, the organisation had. We therefore call its one-shot result on 2021–2025 a post-development evaluation, and claim no pristine blind pass for it.
The evaluation was passed, and is still discounted. One value per parameter does not remove selection at the level above it: the surviving architecture is the survivor of a wider search — factors considered, constructions tried, two whole strategies discarded. That is researcher degrees of freedom, and it is why the live expectation is set below what any backtest shows.
This is also why no compounded backtest return appears on this page. A backtest's job in this process is to reject rules. The record that matters is the live one, published trade by trade, with the date separating backtest from live trading stated in plain sight.
Backtests are graded net of trading friction.
Every simulated trade pays. A per-share commission, plus slippage scaled to each stock's actual liquidity — a thinly traded name pays a realistic spread rather than the free fill of a naive backtest.
Fills happen at the next session's average price. A signal computed after the close cannot trade at that close, so the simulation prices every fill from the following session instead — the average of its high, its low, and double weight on its close, a stand-in for trading spread across the day.
Members see the trades before the modelled fills happen. Reviews run over the weekend and the rebalance email lands before Monday's open; the simulation then fills at that Monday session's average. A member trading during the session faces broadly the conditions the model is charged for, rather than chasing an execution that already happened.
Turnover is a cost even when gross numbers flatter it. Variants that added gross return by trading more were judged on what survived the extra friction, and the live portfolio's bar for a trade is deliberately high — a typical review replaces a name or two of the twenty.
Capacity gets the same honesty. An earlier micro-cap system tested well and was shelved anyway: its names were too thinly traded to absorb meaningful capital, and its drawdowns were too deep to ask anyone else to sit through. A strategy we publish has to be one an ordinary member can actually trade.
Twenty stocks, equal weight, reviewed every four weeks.
The portfolio owns twenty of the highest-ranked investable stocks at equal weight. Twenty is enough that no single mistake is fatal, and few enough that the evidence still shows up in the result. The number was set on principle, then stress-tested around — never chosen by backtest.
Reviews run every four weeks, and a typical one changes a name or two. A holding is sold when its rank has genuinely decayed, never for slipping a point. The gap between the top twenty by rank and the twenty we hold is a deliberate buffer, and it keeps turnover far below what chasing the rank's top twenty would generate.
No margin, no shorts, no hedging overlay, and nothing held back in cash. Long only, fully invested. Every layer removed is a failure mode removed.
Companies in the middle of a takeover are set aside. Once a merger is announced the price belongs to the deal, and the factors have nothing left to say about it.
The benchmark is the S&P 500 Equal Weight.
The ordinary S&P 500 is now roughly a third invested in a handful of enormous technology companies, so beating or losing to it mostly records whether you owned those few. The average stock — equal weight — is the comparison we chose for a twenty-stock book, for that reason.
The S&P 500 is shown beside it anyway, always. Hiding the less flattering line is how a track record becomes marketing.
No single index matches the portfolio's exposures — it tilts smaller than any large-cap benchmark. That is what the factor regression in the testing section is for: exposures are measured directly rather than assumed away by the choice of yardstick.
The Nasdaq 100 plays the adversary. It is the disguise check from the testing section: if the model's results track it too closely, the edge is a growth bet in costume, and the strategy fails regardless of its returns.
What this methodology cannot do.
- It cannot promise the future. Twenty years of evidence bounds our expectations; it guarantees nothing, and every factor in the model has had years of lagging badly.
- It cannot remove judgment. Somebody chose the three-part factor test, the windows, the gates. The choices are documented here precisely because they are choices.
- It cannot see what has not been filed. A fraud, a buyout, a court ruling — nothing in the fundamentals anticipates them. Price usually moves first, and momentum reads that drift, but only once the move is underway; the model still takes the first leg.
- A passed test is not a proof. Our own blind window killed two strategies that had passed everything else. The tests raise the odds; they do not settle them.
The research the model stands on.
The primary literature behind each family, and the replication studies that police the field. A factor entered the model only when its mechanism appears in work like this and survived our own tests; where the replication corpora disagreed about a category, we skipped it.
- Banz (1981). The relationship between return and market value of common stocks. Journal of Financial Economics — size.
- Bernard & Thomas (1989). Post-earnings-announcement drift: delayed price response or risk premium? Journal of Accounting Research — earnings surprise.
- Fama & French (1992). The cross-section of expected stock returns. Journal of Finance — value and size.
- Jegadeesh & Titman (1993). Returns to buying winners and selling losers. Journal of Finance — momentum.
- Chan, Jegadeesh & Lakonishok (1996). Momentum strategies. Journal of Finance — momentum and earnings news.
- Sloan (1996). Do stock prices fully reflect information in accruals and cash flows about future earnings? The Accounting Review — accruals.
- Ang, Hodrick, Xing & Zhang (2006). The cross-section of volatility and expected returns. Journal of Finance — volatility.
- Boudoukh, Michaely, Richardson & Roberts (2007). On the importance of measuring payout yield. Journal of Finance — shareholder yield.
- Cooper, Gulen & Schill (2008). Asset growth and the cross-section of stock returns. Journal of Finance — investment.
- Pontiff & Woodgate (2008). Share issuance and cross-sectional returns. Journal of Finance — issuance.
- Novy-Marx (2013). The other side of value: the gross profitability premium. Journal of Financial Economics — profitability.
- Frazzini & Pedersen (2014). Betting against beta. Journal of Financial Economics — low volatility.
- Harvey, Liu & Zhu (2016). …and the cross-section of expected returns. Review of Financial Studies — multiple-testing discipline.
- Asness, Frazzini & Pedersen (2019). Quality minus junk. Review of Accounting Studies — quality.
- Hou, Xue & Zhang (2020). Replicating anomalies. Review of Financial Studies — replication filter.
- Chen & Zimmermann (2022). Open Source Asset Pricing — the second replication corpus.
- Jensen, Kelly & Pedersen (2023). Is there a replication crisis in finance? Journal of Finance — replication filter.
- Kenneth French's Data Library — the factor return series used in the disguise checks.
The Market Radar puts one of these factors under the lens every Monday, with the week’s best and worst scorer as live examples.
Membership is the Ticker Nerd 20 — the portfolio this methodology runs.