Track Record · Methodology · Updated 2026-08-14

Methodology

What counts as a call, how each one gets graded, and where every derived number comes from. The question this dataset exists to answer is the one readers actually ask: if you had followed this site's calls, would you have made money, and is the author adding information beyond what the market already priced?

What counts as a call

A call is a falsifiable claim published on this site with enough detail to be graded later: a ticker, a stance, a structure, the spot it was struck against (with its session labelled), and the implied move where one was sourced. The trade logs at the bottom of every Options Angle are the raw material; each row there becomes a row here.

Passes are calls. "Do not buy the straddle" is a claim about the world with a P&L consequence for anyone who followed it, and it gets scored as the trade the reader did not make. Three of the costliest calls in the site's July audit were passes.

Nothing here is an executed trade. The author holds no positions; this is a paper record of published calls, disclosed as such everywhere it appears. Returns are what the logged structure would have produced, never what anyone banked.

The two scores: thesis and call result

Every scored row carries two independent grades.

Thesis accuracy asks: did the predicted phenomenon happen? "Realised will beat the 13% implied" is a prediction about volatility, and it is right or wrong regardless of what any trade did.

Call result asks: did the published call win, scored on the whole position? A covered call on a stock that falls 21% collects its premium and still loses a fifth of the capital; that is a loss. A pass on a straddle that would have paid is a loss. A pass on a straddle that would have bled is a win.

Keeping them separate is what makes the ledger diagnostic rather than decorative. Over time it can show, for example, that the site predicts earnings volatility well but expresses it through the wrong structures, or the reverse. On the seeded backlog the two grades coincide on almost every row, because nearly all early calls were passes, where the thesis and the position are the same claim. They diverge as priced structures accumulate.

How calls are graded

  1. Against the settled close, never the first tick. The Super Micro correction is the founding example: a straddle pass graded a win off an early after-hours read of 6-9% became a loss once the stock settled 32.4% higher close-to-close. Extended-hours prints are noted but do not grade a call.
  2. Against the implied move that was logged, not a retrofitted one. Being right on direction and wrong on magnitude is a loss for a premium buyer.
  3. Whole position, not the option leg. Scoring legs in isolation flattered the July record by roughly 15 percentage points of hit rate, which is why the rule exists.
  4. Session labels are part of the record. Every spot says whether it was a close, open, intraday, premarket or after-hours quote. Rows graded off an approximate baseline (the Meta gap) say so in the row.
  5. Ungradeable rows stay visible. A call with no implied logged, no strike pinned, or sources in unresolved conflict is marked unscorable and kept on the ledger. Deleting it would raise the hit rate by erasing sloppiness, which is the opposite of the point.

Categories, and why there is no single score

Every call sits in one bucket: stock direction (buy or avoid the shares), options direction (calls or puts expressing a directional view), volatility (straddles, strangles, premium selling, graded against implied), event (timing calls around a dated catalyst), and macro (sector and index-level calls). Their statistics are reported separately because they are different games: a 60% hit rate selling premium and a 60% hit rate picking direction imply completely different distributions of wins and losses, and averaging them produces a number that describes neither.

Expectancy, priced calls and modelled returns

Expectancy = win rate x average win - loss rate x average loss, computed over priced calls in return-on-risk percent per call. It is the number the hit rate hides: a high hit rate with negative expectancy is a losing strategy that feels good.

A priced call is one with a cost attached. Where a real premium was logged, the return is stated. Where a long straddle was priced only as "roughly the implied move, as a fraction of spot" (the common case in the early logs), the return is modelled: intrinsic value delivered against the premium paid, computed as |realised move| divided by the implied move, minus one, using the top of the quoted implied range so the modelled cost is the conservative one. The model ignores any time value left at exit, which makes it a floor for a winning straddle rather than a flattering estimate. Every derived return in the CSV and JSON carries a return_basis field saying which kind it is.

Passes have no return: no position was proposed, so no P&L exists to assign. The cost of a bad pass is real but it is an opportunity cost, and it is counted where it belongs, in the hit rate and thesis score. Manufacturing a phantom return for the trade nobody proposed would double-count it.

The expectancy cell reads "too few" below five priced calls per bucket. Two priced calls is an anecdote.

Conviction calibration

From 2026-08-14, every new call records a conviction from 1 to 10 at logging time, before the outcome is known, and the validator rejects new calls without one. Once enough convicted calls resolve, the calibration table reports thesis accuracy at each level. The interesting outcomes are the uncomfortable ones: if 8/10 calls hit more often than 10/10 calls, the 10s are overconfidence, and the table will say so.

Calls logged before the cutoff keep a null conviction forever. Assigning confidence to a call whose outcome is already known would corrupt the one table whose entire value is that the number came first.

The benchmark

A call is only worth following if it beats what doing nothing would have done. From 2026-08-14, each new call records the SPY close at entry, and the close at exit when it is graded; the exports then carry the SPY return over the holding period and the call's excess move against it. Seeded rows predate the rule and carry nulls rather than backfilled benchmark prices, for the same reason conviction is not backfilled: the discipline of this dataset is that fields are filled when they were knowable, or never.

Sector benchmarks (SMH for the semiconductor calls, QQQ for large-cap tech) are the intended next step once the SPY column has a real sample.

Open, pending, unscorable

Open: a sourced entry exists and the call has not resolved. Open calls with a live mark show an unrealised figure, kept outside every statistic. Pending: a conditional call whose trigger fired but whose entry price has not been sourced; it carries no invented entry and joins the scored set only when a real one is found. Unscorable: permanently ungradeable, kept visible, excluded from every rate.

Where the ledger starts, and what is missing

The ledger begins with the week of August 3, 2026, transcribed from the site's first published scorecard, which graded 37 calls at a 59% hit rate. Rows carry the scorecard's corrected figures where the scorecard corrected its sources (Super Micro's grade, Cava's grade, Datadog's spot, the Vistra and Oklo spot dates).

Calls made before that week are mostly not reconstructible: the August 1 audit found that of roughly a dozen July pieces carrying options plays, only one published an entry price. The July record survives in summary form (about a 50% hit rate, nearer 40% on whole-position scoring) in the scorecard rules that came out of it, and two July calls that were graded inside later articles (the Micron put-selling suspension, the Micron/SanDisk pair) appear here as unscorable rows. A ledger that pretended to reach further back than its records do would be exactly the kind of track record this page exists not to be.

Corrections

A wrong grade gets corrected in place, with the original grade and the reason recorded in the row's note, the way the Super Micro and Cava rows carry theirs. Corrections change the statistics and the statistics are allowed to get worse; what never happens is a row disappearing. Spot an error, cite the row id from the CSV through the contact form.