The criteria, in detail
The scorecard exists because the options-signal market is packed with figures that few buyers can ever confirm. Three of the five tests carry their own page, because they are the ones services most often fail; the full five-test method lives on the method page.
What the scorecard is actually sorting
A buyer almost never weighs two near-identical services side by side. In practice the choice is between kinds of service — a chat channel, a copy-trading room, a social caller, an aggregator, or an audited desk — and each kind passes or fails the tests as a category, not as a one-off. The scorecard's job is to drag that structural difference into the light, so a glossy presentation cannot hide which category a service belongs to. The table below sorts the field on the two tests that settle most cases: was the call, grade included, sealed before its outcome, and is the win rate shown with its full denominator.
| Service type | Grade sealed first? | Denominator shown? | Why it lands there |
|---|---|---|---|
| Messaging-app channel | No | Rarely | The operator edits and deletes at will; losing calls never appear. |
| Copy-trading room | Rarely | Sometimes | A platform logs results but seldom proves a per-call timestamp. |
| Social-media caller | No | No | Threads are deletable; income often runs through affiliate links. |
| Signal aggregator | No | No | Re-posts others unaudited and inherits every gap in the original. |
| Automated / AI service | Sometimes | Backtest only | A backtest is not a live result, and no person owns the record. |
| Audited, timestamped desk | Yes | Yes | Per-call on-chain receipt with the grade inside it, plus named independent review. |
Only the bottom row answers yes to both, which is the structural case this guide makes for the pick — not that it is louder, but that it belongs to the single category an outsider can check end to end. The three criteria below each take a single test apart in full.
Why these two columns decide most cases
Of the five tests, two do nearly all the work. Sealed before the outcome is the one that cannot be retrofitted: a service either committed its calls and grades in public before they resolved or it did not, and no later polish rewrites that fact. Denominator shown is the one that cannot be faked short of an outright lie: a win rate is evidence only when the full count of calls, losers folded in, sits beside it. A service that clears both has handed you a record you can interrogate. The remaining three tests — the defined-risk shape, the measured grade and clean incentives — are real, but they tend to confirm a verdict the first two have already reached rather than overturn it. That is why the table sorts on the two, and why a service can carry a slick site, a busy community and a confident pitch and still land in a row with two crosses.
The three tests with their own page
Sealed before the outcome
Why a public timestamp - with the grade hashed inside it - is the test a service cannot fake its way around.
A defined-risk structure
What a defined-risk call looks like, and why a bounded worst case is the first thing an options buyer should demand.
Grades that are measured
How an A-to-D conviction grade is calibrated to a model's own returns instead of a mood word.
The remaining two tests — a re-runnable track record and clean (non-affiliate) incentives — are covered on the method page, because they are quicker to check and rarely the deciding factor.