Independent research desk · no sales team

Small caps. Fat tails. A dataset nobody else has.

ScalarDelta is a research engine for professionals who trade the tail of small-cap equities. We built our own corpus of company-level mechanism evidence, trained our own models on it, and score every name on the shape of its forward distribution — not on a price target.

Three tiers, monthly, cancel any time. The call is with a technical trader — we have nobody in sales.

Heavy-tailed versus normal return distribution Two densities with the same peak. The heavy-tailed one puts far more mass in both tails — the losses you must survive and the outcomes that pay for everything else. ScalarDelta scores that shape rather than the average. left tail forward return right tail survive this own this same mean · same peak · different business
Illustrative. The average of a small-cap book is a poor description of it — the tails are the trade, so the tails are what we score.
695
securities in the frozen validation baseline — 95 of them already delisted, kept on purpose
~740
names under daily AI evidence reading, selected by state rather than by size
Walk-forward
on every persisted result: rolling train/test, never one in-sample number
20d
one standard horizon, so two playbooks can actually be compared
What we built

A proprietary dataset, and our own models trained on it

Everyone has prices. Everyone has fundamentals. What almost nobody has is a structured, point-in-time record of why a small company's prospects changed — machine-readable, comparable across thousands of names and years, and built specifically to be tested against forward returns. That is what we spent the last years assembling.

01 / CORPUS

Evidence, not sentiment

Filings, 8-Ks and company news land in our own store every session. Each document is turned into typed mechanism evidence: what changed, in which direction, how material it is, and over what horizon. Not a headline count, not a sentiment score — a structured claim about the business that can be tested against what the stock did next.

02 / MODEL

Trained on a corpus that only exists here

Those evidence vectors cluster into mechanism centroids — the recurring economic storylines behind re-ratings: capacity coming online, a pilot converting to a contract, a decline turning terminal. The clustering, the centroid geometry and the scorers on top are ours, fitted to our corpus. The large language model does one narrow job: reading a filing into structure. The asset is the dataset and the fit.

03 / STATE

Narrative only counts when the state agrees

Every name carries a daily state vector — convexity (BCS, SCS, p_down), compression, position in its 52-week structure, cohort-relative momentum, plus fundamental sleeves for pre-inflection and compounding profiles. A great story in the wrong state is not a trade, and the engine is built to say so.

How a name gets scored

Five steps, run every session

01
Ingest
Prices, fundamentals, filings, insider transactions and the forward event calendar — daily, across tens of thousands of listings.
02
State
Each name is placed in convexity state space: distribution shape, compression, drawdown structure, cohort-relative behaviour.
03
Read
Documents for names whose state qualifies them are read into mechanism evidence and embedded. Quiet names are stored, not read.
04
Fuse
State × mechanism × calendar becomes sleeve scores and playbook candidates, ranked, with an explicit entry status.
05
Validate
Walk-forward and out-of-sample on a frozen universe. What fails gets paused — not renamed and re-marketed.
For the people who will ask hard questions

The alpha logic, stated plainly

If you have built signals yourself, you already know where they break: pooled embeddings, survivorship, a mean that hides the distribution, and a backtest tuned until it is true. Here is how each of those is handled — and where we are still wrong.

A

Per-event matching, never per-security pooling

Mean-pool a company's embeddings into one vector and every cluster collapses to the same point — the signal disappears while the dashboard still looks healthy. Each piece of evidence is matched against the centroids on its own document date. The entry fires on the narrative event, not on a company-level average.

B

We optimise distribution shape, not the mean

Hit rate is the least interesting number we compute. The objective is the shape of the forward distribution: right-tail rate, left-tail exposure, average maximum adverse excursion, net over max drawdown. A 46% hit rate with a fat right tail beats a 62% hit rate that bleeds — which is exactly the trade small-cap specialists are already making by instinct.

C

Anti-survivorship by construction

The validation universe is frozen and includes the dead: 695 securities, 95 of them delisted, deliberately retained. Backtesting on names that are still listed today is backtesting with tomorrow's information. Every persisted result records which universe it ran on and how many names resolved.

D

Statistical gates we will show you

Every playbook carries an effective sample size, a lower confidence bound and a deflated significance measure after multiple testing. Thin results sit in probation instead of in a brochure. On a call we will show you the ones we paused, and why.

E

The signal says look; a rule says size

A worked example from our own research: on volume-backed activations in pressed small caps, 24% of names never reclaim the event-day close and average −13.5% over the next 60 trading days. So activation is a watchlist event, and entry is gated on reclaim. You get the entry status, not just the score.

F

An intraday layer, with its negative results attached

A separately validated reversal engine: a −3% intraday flush, in a name that has not outrun the index by more than 2% over the prior five sessions, with the macro gate evaluated at the fire minute, flat by the close. Cohort-specific and walk-forward tested. We will also tell you what did not survive — the news veto did not, and a −2% trigger is frequency, not edge.

Fit

Built for professionals with a taste for small cap and tail

This is for you if

  • +You run a book in sub-$2bn listings and size for convexity, not for comfort
  • +You want the raw score surface and your own overlay on top of it
  • +You can read a walk-forward table and argue with it
  • +You already know most signals die out-of-sample and want to see the evidence either way
  • +You would rather be handed a state and a rule than a price target

Give this a miss if

  • You want buy and sell alerts to follow without forming a view
  • You need a price target or a position size from us
  • Your mandate is large-cap only — the edge here lives where coverage is thin
  • A cohort where a quarter of names never work is a dealbreaker rather than a known cost
  • You are looking for a newsletter to outsource decisions to
Pricing

Three tiers. Monthly. No sales call required.

Start free and read the output for a few weeks. Plug into what we already compute when you want it in your own stack. Put your own names through the engine when our universe is not your universe.

Signal Digest

See what the engine surfaces before you decide whether to plug into it.

$0 free, forever
  • +One email a week, out before the US open
  • +The ranked shortlist, with each name's sleeve and convexity state
  • +A plain-English mechanism note per name — what changed, and why it matters
  • +Entry status per name: confirmed, reclaimed pullback, waiting, or failed
  • No API access
  • No point-in-time history
  • No nominated tickers
Get the digest

No drip sequence, no upsell mails. One email, unsubscribe by replying STOP.

Most desks start here

Coverage API

Use what is already there: our universe, scored daily, over a REST API.

$999 per month
  • +REST access to the live score surface across our covered universe
  • +Convexity state (BCS, SCS, p_down), sleeve scores, tail score, synthesis legs
  • +Daily point-in-time vintages — backtest against what was actually known that day
  • +Playbook candidates plus persisted backtest history on one 20-day horizon
  • +Mechanism evidence per name, with direction, materiality and horizon
  • +Fair rate limits, no per-call metering, no overage invoice
Request API access

Monthly. Cancel any time. Your overlay, your execution — we do not touch either.

Nominated Coverage

Your names, run through the same engine as ours, under a fair-usage policy.

$4,999 per month
  • +Everything in Coverage API
  • +You nominate the tickers; we ingest, extract, embed and score them as our own
  • +New names enter with their full available filing history, not from day one forward
  • +Your nominations stay yours — scored for you, never published in the digest
  • +A direct line to a technical trader for methodology and model questions
  • +Quarterly walk-forward review of the playbooks that fire on your names
Talk to a trader

Monthly. Fair-usage terms below — they are a compute constraint, stated plainly.

Priced in USD, invoiced monthly, cancel any time. No setup fee, no onboarding fee, no sales commission — there is nobody here to pay one to.

Fair usage policy — Nominated Coverage

Nominated coverage is priced against compute, not against how much you might make from it. So the limits are stated up front rather than buried in a contract.

Nominations
Up to 150 active tickers at a time, swappable once a month. Need more? Say so — it is a compute conversation, not a pricing trick.
Compute
AI evidence reading is state-gated by design. A nominated name that goes quiet is still ingested, but is not re-read until its state moves. That is how the cost stays sane and the corpus stays clean.
Rate & reuse
Generous limits for a human desk and its tooling. Bulk mirroring of the full surface, redistribution or resale is out of scope — ask first, and the answer is usually yes.
Book a meeting

There is nobody here to sell to you

We have no people in sales. Book a call and you get a technical trader: someone who has run the backtests, knows which cohorts the engine is wrong about, and will tell you unprompted when a sample is too thin to believe. Bring your hardest question about the methodology — that is the entire meeting.

Get access

Pick a tier and tell us what you trade

The free digest starts immediately. For the two paid tiers we answer personally, usually within a session or two, because a trader reads every one of these.

  • +No sales sequence, no demo gate, no “book a discovery call”
  • +Tick the box and a technical trader replies, not a rep
  • +Prefer plain email? desk@scalardelta.com

We store your address to answer you and, for the digest, to mail you. No list sharing, no tracking pixels, no drip sequence. Unsubscribe in one click.

Questions we actually get

Before you ask

Do you give buy and sell recommendations?

No. The digest and the API name candidates, the state that surfaced them and an entry status. We never give a target price or a position size. You are the professional; the engine is an instrument, not a manager.

Is this just an LLM wrapper?

No. A commodity language model does one narrow job — turning a filing or a news item into typed, structured evidence. Everything that produces alpha is ours: the corpus, the clustering into mechanism centroids, the convexity scorers, the classifiers and the fusion, all fitted on data we assembled. The dataset and the fit are the product; extraction is plumbing.

Can I see a track record?

On a call, in detail: per-playbook walk-forward and out-of-sample tables, effective sample sizes, drawdowns, and the playbooks we paused. We do not print a headline return number on a marketing page, because a headline return from an in-sample backtest is worth nothing and you know it.

Which markets do you cover?

US listings are deepest. The intraday layer additionally runs live across European and Hong Kong sessions. Nominated tickers can be any listed equity we can source clean data and filings for — if we cannot, we will tell you before you pay for it.

What happens to the tickers I nominate?

They are ingested and scored for you, with their full available history backfilled on entry. They stay out of the public digest. Drop a name and it drops out of active reading.

How is my data handled?

We store your contact details to answer you and, for the digest, to mail you. Your ticker list is used to run coverage for you and nothing else. No list sharing, no resale, no tracking pixels.