Guide

What is OSINT? Open source intelligence explained

A plain definition of OSINT, where the data comes from, how open sources differ from leaked data, and what a legitimate investigation looks like in practice.

Updated 2026-09-257 sections8 min read
01

OSINT meaning, in plain terms

OSINT stands for open source intelligence. It means information gathered from sources that are publicly available, and it is intelligence in the sense of a finished product rather than a secret one. The word open describes the availability of the source, not a licence to do whatever you like with what you find.

The term came out of the military and intelligence world, where it distinguished material drawn from newspapers, broadcasts and public records from material collected by satellites or by human sources. The method has since moved into journalism, threat research, corporate security, fraud investigation and ordinary consumer protection.

What makes work open source is where the material came from, not how clever the analyst is. A search engine result, a company registry, a court filing, a public social profile: all of it is open source material, because anyone who looks can see it.

02

Why the precise meaning matters

OSINT is now used loosely to mean any research about a person or a company, including techniques that are not open at all. Being precise about the term matters, because the legal and ethical position of an investigation depends almost entirely on where the underlying data came from.

If the material sits in a public register, citing it is usually straightforward. If the material came from a leaked database or a compromised account, the situation is entirely different, and calling that OSINT does not change anything.

03

Where open source data actually comes from

Public registries. Company registration, land and property records, court filings, licence registers, and patent or trademark databases. These are published by governments and are usually the most reliable open sources available.

Published media and research. News reporting, press releases, academic papers, annual reports and industry filings.

Open web and technical sources. Websites, documentation, exposed configuration files, certificate transparency logs, DNS records and public code repositories. Much of this is published by accident rather than by intention, which is exactly why it is useful.

Public social and community material. Profiles, posts, public forum threads, review sites, and a public gaming or messaging presence.

What all of these share is that a member of the public can reach them without breaking anything. That is the test that separates an open source from everything else.

04

Open sources versus leaked data

Leaked data is not OSINT. A database taken without permission, or published by someone who was not entitled to publish it, is a different category of material even when copies of it later appear on a public page. Something can be readable and still not be an open source.

The practical consequences follow from that difference. Open source material can generally be cited and reused within the law that applies to you. Leaked material carries obligations, and in many jurisdictions holding or processing it without a lawful basis is itself the problem, however you came by it.

The line matters for the person you are researching too. Reading someone's public filings is ordinary research. Checking whether their address appears in a breach corpus is a different act with a different justification, and it needs one.

Where a platform aggregates breach records, the honest description is that it queries third party sources which hold data collected under those sources' own arrangements, using their published access methods. That is not the same thing as an open source lookup, and a product that blurs the two is not being straight with you.

05

What a legitimate use looks like

A legitimate use starts with a lawful basis and a defined question. You are checking your own exposure, or you hold written authorisation to assess an organisation, or you have a documented reason to investigate a specific identifier that is lawfully in scope.

It continues with data minimisation. Query the identifier the question actually requires, at the point you need it, and collect nothing beyond what the decision needs. A finding that is irrelevant to the question is a liability rather than an asset.

It is honest about uncertainty. Open source findings are evidence rather than proof, and third party records contain errors, duplicates and stale entries. Report what a record says and where it came from, and do not present it as a verified fact about a person.

It respects the limits the law places on decisions. Using a record to assess somebody for credit, insurance, employment, housing or a benefit is regulated in most countries, and the standards those laws require are not met by any breach corpus.

06

What this product does and does not do

OSINT Pro queries third party breach and intelligence sources and returns normalized records relating to an identifier you submit. Every row names the module it came from, so a finding can be traced to its origin. We do not compile or own the underlying datasets, we do not steal data, we do not intrude into private systems, and we do not obtain private material improperly.

We are not a people search site. We do not return home addresses, relatives, contact details scraped from directories, or a profile assembled about a person. We are not a consumer reporting agency, and nothing here may be used for a purpose the Fair Credit Reporting Act regulates.

We do not monitor anything. There is no alerting, no continuous watching of an address or a domain, and no stored search history: an identifier you submit is passed to the sources being queried in order to answer the request and is not written to a record you or we can browse later. What we sell is an on demand lookup, and we would rather describe it that way than imply a capability we have not built.

What we do offer is a straight answer to one narrow question: does this identifier appear in records other people can hold, and if so, which source produced each row. Pricing is published, the first lookup runs in full, and every plan is self serve.

07

Start with a lookup

If your own exposure is the question, the place to start is a lookup on the identifier you care about. Read the source labels first, then the values, and act only on what the records actually contain.

Next step

Check your own identifiers

Run the email address, username or phone number you care about and see which sources hold a record for it.

Run a lookup