Skip to content

Aleph AI Lab

Work in progress

Cyber risk, quantified from what the dark web lets us see

Dated observations going back to 2012 are material no one else holds. Aleph's lab turns it into risk indicators usable by the people whose job is to decide.

What it will bring

A figure that holds up in front of a committee

A risk indicator is only worth something if it survives three questions: where does it come from, how good is it, and what do you do with it. The lab builds for those three.

  • Anticipate instead of observe

    Exposure surfaces on the dark web long before an incident is declared. That lead time is exactly what a risk indicator has to make usable.

  • An indicator you can defend

    In front of a regulator, an actuary or a credit committee, you must be able to explain where a figure comes from. Our work is built for that, not for a demo.

  • Every trade needs a different answer

    An incident response team needs an immediate dated fact. An insurer needs a one-year horizon. A bank thinks in years. The same work serves all three, each in its own language.

  • Since 2012, not a snapshot

    A point-in-time scan only sees a state. Archive depth tells you when an exposure began, how it evolved, and whether it is still dormant somewhere.

The first project

Millions of images no index knows how to read

A considerable share of what we collect on the dark and deep webs is not text: screenshots, photographs of documents, scans of identity papers, statements. The critical information is in there, and no full-text engine sees it. We train models to recognise the nature of these files, to make that share of exposure searchable.

What comes in

Images, with no text describing them

What comes out

A recognised nature, a qualified exposure

  • Internal document0
  • Invoice, statement, purchase order0
  • Screenshot of a system0
  • Identity document0
  • Payment instrument0
  • Site plan or diagram0

An image is not text: what it shows appears in no full-text index. Automatically recognising the nature of what circulates as images makes searchable a share of exposure that today escapes every engine. It is the lab's first project, and it runs on millions of files already collected.

Schematic. The natures illustrate the typing work; the proportions are not a measurement made on the index.

The method, in five steps

From the file to the score

A risk score is not read off a dashboard: it is estimated on a file built for that purpose, then submitted to checks that are entitled to fail it. Here is the full sequence, from the first byte accumulated to the figure handed to the client.

  1. 1 Accumulation
  2. 2 Typing
  3. 3 Splitting
  4. 4 Checks
  5. 5 Score

The entity-period file

One row per company and per month

0

rows

300 000
entities
120
months
8 – 12
retained variables

Each row carries the company's state as it was visible that month, rebuilt in replay — never copied from today.

Internal methodology paper, August 2026. The precision quoted comes from a Bayesian calculation on the incident base rates published by Cyentia (IRIS 2020).

The lock

The hard part is not the model, it is the label

Everyone knows how to fit a regression. Almost nobody has a dated, public ground truth independent of the engine that produces the variables. That separation decides a score's worth, and it is settled before the first line of code.

Aleph index · the variables

  • Dated appearances, archived since 2012
  • Channel: forum, paste site, marketplace
  • The entity's exposure footprint
  • First indexing date, never rewritten

Leak sites · the target

  • Name of the victim entity
  • First publication date
  • Claiming group

Entity-period file

One row per month, variables rebuilt as at date t

The rule that is not negotiable

The target must come from a sensor strictly independent of the engine. If the index serves as both variable and label, the score no longer measures risk: it measures the index's internal consistency. That is grounds for immediate rejection, not a flaw to patch.

Internal methodology paper, August 2026. Timestamped, immutable index; target built from the public extortion-site aggregators.

The result that governs everything

The same model, one hundred and fifty-four times more useful

A model's performance says nothing about its usefulness. What decides is the rarity of the event over the requested window: with detection rate and false-alert rate unchanged, a score's precision varies by a factor of 154 between seven days and seven years. That arithmetic cannot be dodged.

Precision: share of flagged entities that really suffer an incident

  • Response team · mid-cap

    7 days

    %

    251 false alerts for one true

  • Insurer · SME

    12 months

    %

    29 false alerts for one true

  • Insurer · mid-cap

    12 months

    %
  • Insurer · large account

    12 months

    %
  • Bank · SME

    7 cumulated years

    %
  • Bank · mid-cap

    7 cumulated years

    %

    fewer than two false alerts for three true

precision = p · TPR ⁄ ( p · TPR + (1 − p) · FPR )

p is the base rate over the window, TPR the detection rate (90 %), FPR the false-alert rate (10 %). From one line to the next, only p changes.

A factor of 154 separates the two extremes, with a strictly identical model. That result decides the shape of the products: a response team must not receive a score, because at seven days even a model excellent on paper would send it two hundred and fifty-one false alerts for one true. What it must receive is a dated fact.

Bayesian calculation on the incident base rates published by Cyentia (IRIS 2020), cumulated over the window, and on the best published performance in the field (Liu et al., USENIX Security 2015).

The chosen form

One scorecard, three products

The lab does not look for the most sophisticated model, but for the one that survives a rare event and can be explained. A single estimation comes out of it, and three products that nest into one another, each in its trade's unit of time.

  1. 01

    Discretisation

    Each variable is cut into four to six ordered classes. A risk officer reads “between six and twelve months of exposure”, not a seven-decimal coefficient.

  2. 02

    Regression

    Eight to twelve variables, and a complementary log-log link suited to rare events. The discrete-time survival likelihood factorises exactly: it is an ordinary logistic regression.

  3. 03

    Calibration

    Three successive corrections: the rarity of the event, anchoring on base rates by revenue band, then the known gap between incidents that occur and incidents that get reported.

Pr( s < T ≤ s + w | T > s ) = 1 − ∏ ( 1 − h(u) )

h(u) is the hazard of occurrence in month u for a company that has suffered nothing so far. w is one week, twelve months or eighty-four months. One coefficient table, three products that nest into one another.

  1. in days

    The trigger

    A published decision tree, resting on a dated fact. No probability: at that scale it would mislead.

  2. in months

    The score

    Probability, rank within the portfolio and points, side by side. Three readings of the same calculation, for three audiences.

  3. in years

    The term structure

    Derived from the twelve months by composition, never estimated directly: nobody has seven years of clean labels.

Why a scorecard and not a black box

On a rare event, variance dominates bias: a tree model trained on a few hundred events mostly learns noise, and proves it by losing most of its performance on months it has never seen.

Internal methodology paper, August 2026. The three products come out of the same estimation: only the window w changes.

Three trades

The same material, three ways to use it

The lab works with teams that actually decide, not with textbook cases. That is what shapes the final indicators.

  1. 01Response teams

    A fact, right now

    This identifier, on this domain, appeared on this date. Something you can act on within the hour.

  2. 02Insurers

    A one-year horizon

    Enough to price, to prioritise a portfolio review and to justify an underwriting decision.

  3. 03Banks & lenders

    A multi-year view

    A counterparty's cyber risk over the life of a commitment, in a format validation teams know how to review.

  4. 04You

    Lab partner

    The work is calibrated with field teams. It is open, and now is the best moment to join.

Take part in the work

Insurers, banks, response teams: the lab is looking for calibration partners, join us.

Talk to the lab