Aleph AI Lab
Work in progressCyber risk, quantified from what the dark web lets us see
Dated observations going back to 2012 are material no one else holds. Aleph's lab turns it into risk indicators usable by the people whose job is to decide.
0,41 %
12-month probability
Factors retained
What it will bring
A figure that holds up in front of a committee
A risk indicator is only worth something if it survives three questions: where does it come from, how good is it, and what do you do with it. The lab builds for those three.
Anticipate instead of observe
Exposure surfaces on the dark web long before an incident is declared. That lead time is exactly what a risk indicator has to make usable.
An indicator you can defend
In front of a regulator, an actuary or a credit committee, you must be able to explain where a figure comes from. Our work is built for that, not for a demo.
Every trade needs a different answer
An incident response team needs an immediate dated fact. An insurer needs a one-year horizon. A bank thinks in years. The same work serves all three, each in its own language.
Since 2012, not a snapshot
A point-in-time scan only sees a state. Archive depth tells you when an exposure began, how it evolved, and whether it is still dormant somewhere.
The first project
Millions of images no index knows how to read
A considerable share of what we collect on the dark and deep webs is not text: screenshots, photographs of documents, scans of identity papers, statements. The critical information is in there, and no full-text engine sees it. We train models to recognise the nature of these files, to make that share of exposure searchable.
What comes in
Images, with no text describing them
What comes out
A recognised nature, a qualified exposure
- Internal document0
- Invoice, statement, purchase order0
- Screenshot of a system0
- Identity document0
- Payment instrument0
- Site plan or diagram0
An image is not text: what it shows appears in no full-text index. Automatically recognising the nature of what circulates as images makes searchable a share of exposure that today escapes every engine. It is the lab's first project, and it runs on millions of files already collected.
Schematic. The natures illustrate the typing work; the proportions are not a measurement made on the index.
The method, in five steps
From the file to the score
A risk score is not read off a dashboard: it is estimated on a file built for that purpose, then submitted to checks that are entitled to fail it. Here is the full sequence, from the first byte accumulated to the figure handed to the client.
- 1 · Accumulation Accumulation
- 2 · Typing Typing
- 3 · Splitting Splitting
- 4 · Checks Checks
- 5 · Score Score
The entity-period file
One row per company and per month
0
rows
- 300 000
- entities
- 120
- months
- 8 – 12
- retained variables
Each row carries the company's state as it was visible that month, rebuilt in replay — never copied from today.
Internal methodology paper, August 2026. The precision quoted comes from a Bayesian calculation on the incident base rates published by Cyentia (IRIS 2020).
The lock
The hard part is not the model, it is the label
Everyone knows how to fit a regression. Almost nobody has a dated, public ground truth independent of the engine that produces the variables. That separation decides a score's worth, and it is settled before the first line of code.
Aleph index · the variables
- Dated appearances, archived since 2012
- Channel: forum, paste site, marketplace
- The entity's exposure footprint
- First indexing date, never rewritten
Leak sites · the target
- Name of the victim entity
- First publication date
- Claiming group
Entity-period file
One row per month, variables rebuilt as at date t
The rule that is not negotiable
The target must come from a sensor strictly independent of the engine. If the index serves as both variable and label, the score no longer measures risk: it measures the index's internal consistency. That is grounds for immediate rejection, not a flaw to patch.
Internal methodology paper, August 2026. Timestamped, immutable index; target built from the public extortion-site aggregators.
The result that governs everything
The same model, one hundred and fifty-four times more useful
A model's performance says nothing about its usefulness. What decides is the rarity of the event over the requested window: with detection rate and false-alert rate unchanged, a score's precision varies by a factor of 154 between seven days and seven years. That arithmetic cannot be dodged.
Precision: share of flagged entities that really suffer an incident
Response team · mid-cap
7 days
— %251 false alerts for one true
Insurer · SME
12 months
— %29 false alerts for one true
Insurer · mid-cap
12 months
— %Insurer · large account
12 months
— %Bank · SME
7 cumulated years
— %Bank · mid-cap
7 cumulated years
— %fewer than two false alerts for three true
precision = p · TPR ⁄ ( p · TPR + (1 − p) · FPR )
p is the base rate over the window, TPR the detection rate (90 %), FPR the false-alert rate (10 %). From one line to the next, only p changes.
A factor of 154 separates the two extremes, with a strictly identical model. That result decides the shape of the products: a response team must not receive a score, because at seven days even a model excellent on paper would send it two hundred and fifty-one false alerts for one true. What it must receive is a dated fact.
Bayesian calculation on the incident base rates published by Cyentia (IRIS 2020), cumulated over the window, and on the best published performance in the field (Liu et al., USENIX Security 2015).
The chosen form
One scorecard, three products
The lab does not look for the most sophisticated model, but for the one that survives a rare event and can be explained. A single estimation comes out of it, and three products that nest into one another, each in its trade's unit of time.
- 01
Discretisation
Each variable is cut into four to six ordered classes. A risk officer reads “between six and twelve months of exposure”, not a seven-decimal coefficient.
- 02
Regression
Eight to twelve variables, and a complementary log-log link suited to rare events. The discrete-time survival likelihood factorises exactly: it is an ordinary logistic regression.
- 03
Calibration
Three successive corrections: the rarity of the event, anchoring on base rates by revenue band, then the known gap between incidents that occur and incidents that get reported.
Pr( s < T ≤ s + w | T > s ) = 1 − ∏ ( 1 − h(u) )
h(u) is the hazard of occurrence in month u for a company that has suffered nothing so far. w is one week, twelve months or eighty-four months. One coefficient table, three products that nest into one another.
in days
The trigger
A published decision tree, resting on a dated fact. No probability: at that scale it would mislead.
in months
The score
Probability, rank within the portfolio and points, side by side. Three readings of the same calculation, for three audiences.
in years
The term structure
Derived from the twelve months by composition, never estimated directly: nobody has seven years of clean labels.
Why a scorecard and not a black box
On a rare event, variance dominates bias: a tree model trained on a few hundred events mostly learns noise, and proves it by losing most of its performance on months it has never seen.
Internal methodology paper, August 2026. The three products come out of the same estimation: only the window w changes.
Three trades
The same material, three ways to use it
The lab works with teams that actually decide, not with textbook cases. That is what shapes the final indicators.
01Response teams
A fact, right now
This identifier, on this domain, appeared on this date. Something you can act on within the hour.
02Insurers
A one-year horizon
Enough to price, to prioritise a portfolio review and to justify an underwriting decision.
03Banks & lenders
A multi-year view
A counterparty's cyber risk over the life of a commitment, in a format validation teams know how to review.
04You
Lab partner
The work is calibrated with field teams. It is open, and now is the best moment to join.
Take part in the work
Insurers, banks, response teams: the lab is looking for calibration partners, join us.