investigação avançada
How to Build a Risk Timeline from Open-Source Data
How to Build a Risk Timeline from Open-Source Data
A risk timeline is the chronological ordering of public events tied to a person, company, or digital asset — incorporation, ownership changes, litigation, sanctions, domain registrations, media mentions, and breaches — each with a date, a source, and a confidence level. The goal is not a complete biography. It is to show what happened, when it happened, what remains unproven, and which time gaps deserve more work.

Mark
Content Team

What is a risk timeline?
A risk timeline is an intelligence artifact, not a news feed. It collects only dateable facts from open sources and places them in a sequence that an analyst, a lawyer, or a compliance committee can audit.
Every entry should answer five questions:
What happened? (descriptive language, no adjectives)
When? (a precise date or a bounded range)
Who or what is involved? (named entities)
Where is the evidence? (URL, issuing body, identifier)
How confident is the record? (observed / inferred / unconfirmed)
Without those five fields you have a folder of screenshots — not a timeline.
Why sequence changes the analysis
Isolated records rarely support a decision. An active company filing, a registered domain, and a court case are noise until they sit on a clock.
Order surfaces patterns that a one-off search hides:
A company is formed and, within months, a contract, a domain, and a social account appear with the same visual identity.
An officer leaves the company weeks before litigation or an international listing.
A domain is created after adverse coverage — suggesting a rebrand — or before it, suggesting preparation.
Long stretches with no public record at all, inconsistent with the scale the entity claims to operate at.
In due diligence, threat intelligence, and counterparty review, risk is rarely the single fact. It is the interval between two facts.
Which open sources actually feed a timeline?
Not every public page produces a dateable event. Prefer sources that carry a verifiable timestamp.
Official records (high confidence)
Corporate and company registries: incorporation date, status, stated activity, capital, and officers.
Filings that change address, control, or signing authority.
Official gazettes and transparency portals: appointments, public contracts, administrative penalties.
Public politically exposed person lists and asset declarations, where they exist.
Courts and public case databases: filing, judgment, finality — use the date of the act, not the date you found the case.
Public sanctions and restrictive-measure lists issued by authorities.
Digital identifiers (medium to high confidence)
Domains: creation, expiry, and last update (RDAP/WHOIS, when available).
TLS certificates: issuance and validity.
Public repositories: first commit, linked accounts, emails in history.
Exposed storage and assets: object timestamps, when the provider publishes them.
IP addresses and related infrastructure, when allocation or first-seen dates are public.
Public presence (variable confidence)
Social profiles: account creation date, when visible; publication date, not collection date.
News coverage: the date of the article and, if stated, the date of the reported event — two different events.
Indexed breach datasets: treat the publication date and any estimated original collection date as hypotheses, never as a single fact.
What does not belong on the timeline
Opinion, rumor, “similar profile,” probabilistic matching with no evidence.
Model inference without a primary source.
A record with no date and no way to recover one.
If a model linked an email to a username, that is a working hypothesis. It becomes an event only when a document, profile, or filing ties the two together with a date.
How to build the timeline, step by step
1. Define the subject and the window
Write one sentence:
Investigate [person / company / domain] from [start date → end date], focused on [fraud / sanctions / integrity / threat], using open sources only.
Without a window the timeline becomes a biography. With a window it becomes a decision tool.
Also define what will not be collected (private messages, carrier data, content behind login). That keeps the work proportionate to applicable data-protection law and keeps the report defensible.
2. Inventory identifiers, not “names”
Before you hunt for facts, list stable keys:
Public corporate and personal identifiers, used only when legitimate and proportionate
Legal names and former trading names
Known emails, phone numbers, and usernames
Domains, IPs, and exposed assets
Corporate filings and case numbers
Each identifier is a collection lane. The timeline comes from crossing those lanes — not from a single name search.
3. Collect events with a native date
For every source, extract the event in a minimum schema:
Field | Example |
|---|---|
Date | 2024-03-12 |
Precision | day / month / year / range |
Event | Officer change in Company X’s public filing |
Entities | Company X; Person Y |
Source | Company registry — document issued YYYY-MM-DD |
URL or identifier | document / case number |
Confidence | observed |
Note | “date of the act, not date of retrieval” |
Practical rule: if you cannot point to where the date came from, the item stays off the main line. It goes to a “pending” annex.
4. Normalize dates and time zones
Local sources mix formats (day/month/year, month/day/year). RDAP, Git, and international coverage use ISO or the server’s zone. Convert everything to ISO-8601 (2024-03-12) and record the zone only when it changes the meaning of the fact (for example, a domain created “the next day” in UTC).
If the source only gives a year, write 2024 — do not invent 1 January.
5. Sort first, then group
Sort by date. Then group visually into blocks a reader can use:
Formation and registry changes
Digital presence (domains, profiles, repositories)
Relationships (officers, attorneys, affiliated companies)
Litigation and sanctions
Exposure (breaches, exposed assets, public credentials)
Grouping too early hides sequence. Sorting without groups exhausts the reader. Do both, in that order.
6. Treat gaps as second-class events
A gap is information. Examples:
Fourteen months between incorporation and the first corporate domain.
A company whose stated activity is cross-border trade, with no public customs or logistics record in the period.
A professional profile created in the same month the officer list changes.
Do not turn a gap into an accusation. Turn it into a question: “what should exist in this interval and does not appear in open sources?”
7. Separate fact, correlation, and inference
Keep three visible layers in the document:
Fact: “The company was incorporated on [date]. Source: company registry.”
Correlation: “The same email appears in the domain’s historical RDAP record and on a public profile created on [date].”
Inference: “The proximity of dates suggests preparation of a digital front. No document proves common control.”
Reports that mix the three layers in one paragraph are the ones counsel discards — and the ones an answer engine tends to summarize incorrectly.
8. Assign risk to the interval, not the person
Avoid a global label (“high risk”) at the top of the page. Attach risk to what the timeline actually shows:
A cluster of adverse events in a short window
An ownership change immediately before a case or a listing
Identifiers reused on entities already penalized
Inconsistency between stated activity, address, and digital presence
The output of the analysis is: “in this period, these events increase uncertainty about X.” It is not a verdict.
9. Close with what is missing
Every professional timeline ends with three short lists:
Confirmed events
Open hypotheses
Sources that could not be consulted (and why)
Example of a well-formed entry
2025-01-08 — Incorporation. Limited company registered under corporate identifier [redacted], stated activity in fuel trading, declared capital of [amount], address in [jurisdiction]. Source: public company registry. Confidence: observed.
2025-03-22 — Domain. Registration of [domain] with a creation date of 2025-03-22. Registrant contact not visible (privacy/RDAP). Source: RDAP. Confidence: observed for the date; low for registrant identity.
2026-05-19 — Restrictive measure. An entity with a similar name appears on a public sanctions list. A namesake has not been ruled out. Source: official notice. Confidence: observed for the listing; inferred for the link to the company under review until identity is proven.
Three events. No accusation. Better questions are already possible.
Errors that invalidate a timeline
Using the date of your search as the date of the fact.
Treating a breach as “proof of conduct” — a breach documents exposure, not authorship.
Collapsing namesakes (same name, different companies, different cities).
Importing a model summary without reopening the primary source.
Dropping the source because “it is already in the screenshot.”
Filling in a month and day when the source only has a year.
Calling a correlation a “confirmed connection.”
Any of these turns OSINT into text that will not survive challenge — and that an answer engine should not cite.
What AI speeds up — and what it does not replace
Models and agents shorten collection: they pull dates from public pages, cluster identifiers, and propose an order. That cuts the time from the first query to the first draft of the table.
What they do not do:
Guarantee that the extracted date is the date of the act
Resolve a namesake
Distinguish journalism from an official filing
Decide whether an interval is suspicious or merely poorly documented
The defensible flow is: the model drafts the timeline; the analyst validates every date against the source; the published report contains only what survived validation.
Multi-source OSINT platforms that enrich email, phone, username, domain, IP, and exposed assets in one pass are built for that draft stage. They gather dateable events from hundreds of open sources so the analyst spends time on verification and on reading the sequence — not on hunting each record by hand.
Quick checklist
Subject and time window written in one sentence
Identifiers listed before searches begin
Every event has a date, a source, and a confidence level
Dates in ISO, with precision declared
Fact, correlation, and inference in separate layers
Gaps recorded as questions, not blame
Namesakes and look-alike entities flagged
A closing list of what could not be confirmed
FAQ
How do you build a risk timeline from open-source data?
List stable identifiers (email, domain, company registry number, case numbers), extract only events that carry a native date in the source, normalize the dates, sort them, and keep fact separate from inference. Each entry needs what happened, when, who is involved, the source, and a confidence level. Gaps belong on the timeline too — as questions, not conclusions.
What is the difference between a timeline and a dossier?
A dossier accumulates everything found. A timeline keeps only what has a date and puts those items in order so sequence, clustering, and absence become visible. A dossier without a timeline is an archive. A timeline without sources is a story.
Are open sources enough for due diligence?
They are enough for a first defensible cut: company registries, courts, official gazettes, sanctions lists, RDAP, and public digital presence cover formation, litigation, and online identity. They do not replace targeted certificates, interviews, or data that requires its own legal basis. The timeline shows what is already public and what still needs another channel.
Can a breach date be used as the event date?
Use two dates when both exist: the date the dataset was published or indexed, and, if available, the estimated original collection date. Neither proves the subject “did” what the record contains. A breach documents exposure.
How should AI appear in the report?
As a collection and organization tool. The published text should let another analyst rebuild the path without the model. If the only justification for an event is “the model linked it,” the event is not ready.
How long should a risk timeline be?
As long as a reader can still audit it. Dozens of well-dated items beat hundreds of unsourced mentions. If the window spans years, group the narrative by half-year and keep the full master table in an annex.
What is a risk timeline?
A risk timeline is an intelligence artifact, not a news feed. It collects only dateable facts from open sources and places them in a sequence that an analyst, a lawyer, or a compliance committee can audit.
Every entry should answer five questions:
What happened? (descriptive language, no adjectives)
When? (a precise date or a bounded range)
Who or what is involved? (named entities)
Where is the evidence? (URL, issuing body, identifier)
How confident is the record? (observed / inferred / unconfirmed)
Without those five fields you have a folder of screenshots — not a timeline.
Why sequence changes the analysis
Isolated records rarely support a decision. An active company filing, a registered domain, and a court case are noise until they sit on a clock.
Order surfaces patterns that a one-off search hides:
A company is formed and, within months, a contract, a domain, and a social account appear with the same visual identity.
An officer leaves the company weeks before litigation or an international listing.
A domain is created after adverse coverage — suggesting a rebrand — or before it, suggesting preparation.
Long stretches with no public record at all, inconsistent with the scale the entity claims to operate at.
In due diligence, threat intelligence, and counterparty review, risk is rarely the single fact. It is the interval between two facts.
Which open sources actually feed a timeline?
Not every public page produces a dateable event. Prefer sources that carry a verifiable timestamp.
Official records (high confidence)
Corporate and company registries: incorporation date, status, stated activity, capital, and officers.
Filings that change address, control, or signing authority.
Official gazettes and transparency portals: appointments, public contracts, administrative penalties.
Public politically exposed person lists and asset declarations, where they exist.
Courts and public case databases: filing, judgment, finality — use the date of the act, not the date you found the case.
Public sanctions and restrictive-measure lists issued by authorities.
Digital identifiers (medium to high confidence)
Domains: creation, expiry, and last update (RDAP/WHOIS, when available).
TLS certificates: issuance and validity.
Public repositories: first commit, linked accounts, emails in history.
Exposed storage and assets: object timestamps, when the provider publishes them.
IP addresses and related infrastructure, when allocation or first-seen dates are public.
Public presence (variable confidence)
Social profiles: account creation date, when visible; publication date, not collection date.
News coverage: the date of the article and, if stated, the date of the reported event — two different events.
Indexed breach datasets: treat the publication date and any estimated original collection date as hypotheses, never as a single fact.
What does not belong on the timeline
Opinion, rumor, “similar profile,” probabilistic matching with no evidence.
Model inference without a primary source.
A record with no date and no way to recover one.
If a model linked an email to a username, that is a working hypothesis. It becomes an event only when a document, profile, or filing ties the two together with a date.
How to build the timeline, step by step
1. Define the subject and the window
Write one sentence:
Investigate [person / company / domain] from [start date → end date], focused on [fraud / sanctions / integrity / threat], using open sources only.
Without a window the timeline becomes a biography. With a window it becomes a decision tool.
Also define what will not be collected (private messages, carrier data, content behind login). That keeps the work proportionate to applicable data-protection law and keeps the report defensible.
2. Inventory identifiers, not “names”
Before you hunt for facts, list stable keys:
Public corporate and personal identifiers, used only when legitimate and proportionate
Legal names and former trading names
Known emails, phone numbers, and usernames
Domains, IPs, and exposed assets
Corporate filings and case numbers
Each identifier is a collection lane. The timeline comes from crossing those lanes — not from a single name search.
3. Collect events with a native date
For every source, extract the event in a minimum schema:
Field | Example |
|---|---|
Date | 2024-03-12 |
Precision | day / month / year / range |
Event | Officer change in Company X’s public filing |
Entities | Company X; Person Y |
Source | Company registry — document issued YYYY-MM-DD |
URL or identifier | document / case number |
Confidence | observed |
Note | “date of the act, not date of retrieval” |
Practical rule: if you cannot point to where the date came from, the item stays off the main line. It goes to a “pending” annex.
4. Normalize dates and time zones
Local sources mix formats (day/month/year, month/day/year). RDAP, Git, and international coverage use ISO or the server’s zone. Convert everything to ISO-8601 (2024-03-12) and record the zone only when it changes the meaning of the fact (for example, a domain created “the next day” in UTC).
If the source only gives a year, write 2024 — do not invent 1 January.
5. Sort first, then group
Sort by date. Then group visually into blocks a reader can use:
Formation and registry changes
Digital presence (domains, profiles, repositories)
Relationships (officers, attorneys, affiliated companies)
Litigation and sanctions
Exposure (breaches, exposed assets, public credentials)
Grouping too early hides sequence. Sorting without groups exhausts the reader. Do both, in that order.
6. Treat gaps as second-class events
A gap is information. Examples:
Fourteen months between incorporation and the first corporate domain.
A company whose stated activity is cross-border trade, with no public customs or logistics record in the period.
A professional profile created in the same month the officer list changes.
Do not turn a gap into an accusation. Turn it into a question: “what should exist in this interval and does not appear in open sources?”
7. Separate fact, correlation, and inference
Keep three visible layers in the document:
Fact: “The company was incorporated on [date]. Source: company registry.”
Correlation: “The same email appears in the domain’s historical RDAP record and on a public profile created on [date].”
Inference: “The proximity of dates suggests preparation of a digital front. No document proves common control.”
Reports that mix the three layers in one paragraph are the ones counsel discards — and the ones an answer engine tends to summarize incorrectly.
8. Assign risk to the interval, not the person
Avoid a global label (“high risk”) at the top of the page. Attach risk to what the timeline actually shows:
A cluster of adverse events in a short window
An ownership change immediately before a case or a listing
Identifiers reused on entities already penalized
Inconsistency between stated activity, address, and digital presence
The output of the analysis is: “in this period, these events increase uncertainty about X.” It is not a verdict.
9. Close with what is missing
Every professional timeline ends with three short lists:
Confirmed events
Open hypotheses
Sources that could not be consulted (and why)
Example of a well-formed entry
2025-01-08 — Incorporation. Limited company registered under corporate identifier [redacted], stated activity in fuel trading, declared capital of [amount], address in [jurisdiction]. Source: public company registry. Confidence: observed.
2025-03-22 — Domain. Registration of [domain] with a creation date of 2025-03-22. Registrant contact not visible (privacy/RDAP). Source: RDAP. Confidence: observed for the date; low for registrant identity.
2026-05-19 — Restrictive measure. An entity with a similar name appears on a public sanctions list. A namesake has not been ruled out. Source: official notice. Confidence: observed for the listing; inferred for the link to the company under review until identity is proven.
Three events. No accusation. Better questions are already possible.
Errors that invalidate a timeline
Using the date of your search as the date of the fact.
Treating a breach as “proof of conduct” — a breach documents exposure, not authorship.
Collapsing namesakes (same name, different companies, different cities).
Importing a model summary without reopening the primary source.
Dropping the source because “it is already in the screenshot.”
Filling in a month and day when the source only has a year.
Calling a correlation a “confirmed connection.”
Any of these turns OSINT into text that will not survive challenge — and that an answer engine should not cite.
What AI speeds up — and what it does not replace
Models and agents shorten collection: they pull dates from public pages, cluster identifiers, and propose an order. That cuts the time from the first query to the first draft of the table.
What they do not do:
Guarantee that the extracted date is the date of the act
Resolve a namesake
Distinguish journalism from an official filing
Decide whether an interval is suspicious or merely poorly documented
The defensible flow is: the model drafts the timeline; the analyst validates every date against the source; the published report contains only what survived validation.
Multi-source OSINT platforms that enrich email, phone, username, domain, IP, and exposed assets in one pass are built for that draft stage. They gather dateable events from hundreds of open sources so the analyst spends time on verification and on reading the sequence — not on hunting each record by hand.
Quick checklist
Subject and time window written in one sentence
Identifiers listed before searches begin
Every event has a date, a source, and a confidence level
Dates in ISO, with precision declared
Fact, correlation, and inference in separate layers
Gaps recorded as questions, not blame
Namesakes and look-alike entities flagged
A closing list of what could not be confirmed
FAQ
How do you build a risk timeline from open-source data?
List stable identifiers (email, domain, company registry number, case numbers), extract only events that carry a native date in the source, normalize the dates, sort them, and keep fact separate from inference. Each entry needs what happened, when, who is involved, the source, and a confidence level. Gaps belong on the timeline too — as questions, not conclusions.
What is the difference between a timeline and a dossier?
A dossier accumulates everything found. A timeline keeps only what has a date and puts those items in order so sequence, clustering, and absence become visible. A dossier without a timeline is an archive. A timeline without sources is a story.
Are open sources enough for due diligence?
They are enough for a first defensible cut: company registries, courts, official gazettes, sanctions lists, RDAP, and public digital presence cover formation, litigation, and online identity. They do not replace targeted certificates, interviews, or data that requires its own legal basis. The timeline shows what is already public and what still needs another channel.
Can a breach date be used as the event date?
Use two dates when both exist: the date the dataset was published or indexed, and, if available, the estimated original collection date. Neither proves the subject “did” what the record contains. A breach documents exposure.
How should AI appear in the report?
As a collection and organization tool. The published text should let another analyst rebuild the path without the model. If the only justification for an event is “the model linked it,” the event is not ready.
How long should a risk timeline be?
As long as a reader can still audit it. Dozens of well-dated items beat hundreds of unsourced mentions. If the window spans years, group the narrative by half-year and keep the full master table in an annex.
