Grey web

Grey web intelligence.

Collecting public content before it disappears, and keeping it usable afterwards. The part of OSINT that stops assuming the source will still be there when you go back.

Short answer

Grey web intelligence is the discipline of collecting, preserving and analysing content that was published publicly and has since fallen out of public view. What separates it from conventional OSINT is one method: collection happens at publication time rather than at query time, so a deletion becomes a dated event instead of an empty result set. The capability rests on elapsed time, continuous ingestion, retention through the deletion, identity resolution and a defensible legal footing.
01

The definition

Grey web intelligence is the discipline of collecting, preserving and analysing content that was published publicly and has since fallen out of public view. The grey web is what a post becomes after its author deletes it, a moderator removes it, the community goes private or is banned, or the platform closes the API that made it enumerable.

One method separates the discipline from ordinary OSINT: collection happens at publication time rather than at query time. A live search returns what survives today and says nothing about what is missing. An archive built continuously can answer the harder question, which is what was there and when it stopped being there.

02

Why live collection cannot do this

Every tool that queries a platform on demand inherits the platform's present tense. Ask it about an account that scrubbed itself last month and it returns an empty result set, which reads exactly like an account that never posted. The absence is not flagged, because the tool has no basis for knowing anything was there.

That failure mode is quiet and it is common. Content that draws attention is content that gets deleted, so the material most likely to matter to an investigation is the material most likely to be gone by the time anyone looks. Recruitment posts, tooling discussion, bragging, the message that was up for four hours before someone thought better of it.

It is also measurable, which is unusual for a negative. Re-checking a sample of Reddit comments a week after they were captured tells you what fraction has left public view in that window. Our first runs put it in the low single digits per week, weighted towards moderator suppression rather than author deletion. Small samples so far, and we will publish the method and the interval rather than a round number.

03

What a grey web capability actually requires

Five things, and four of them cannot be bought late.

Elapsed time. An archive that started in 2015 holds material no budget can buy in 2026. This is the entire barrier to entry in the field and the reason it has so few real participants.

Continuous ingestion, not crawling on demand. The system has to be reading the platform when the content is published, not when a customer asks.

Retention through the deletion event. Keeping the capture after the original is gone, with the disappearance recorded as a dated fact rather than a row that quietly vanishes.

Identity resolution over time. Accounts rename, die and get remade. A record that cannot connect a persona to its earlier self is a pile of strings.

A defensible legal and ethical footing. An archive of deleted public speech is a serious thing to hold. It needs a stated lawful basis, restricted access, and a real route for a person to reduce their public exposure.

04

Where the market actually sits

Three categories of vendor get mentioned in this conversation, and none of them is doing grey web intelligence, which is worth saying plainly rather than treating as a slogan.

Access vendors such as Bright Data, Oxylabs, Zyte and Apify sell the ability to fetch a page: proxies, unblocking, scraper APIs, priced per request or per gigabyte. They are good at what they do. They hold no archive, so they can retrieve only what is live at the moment you ask.

Listening and monitoring platforms such as Brandwatch, Talkwalker, Meltwater and Exorde sell sentiment, mentions and dashboards over current content. Their unit of value is the trend. Nothing they hold survives a deletion, because the trend does not need it to.

Investigation platforms such as Babel Street, ShadowDragon, Skopenow, Maltego, Flashpoint and Recorded Future serve the same buyers we do, but they are analyst workflow products layered on collection they mostly licence in. They compete on breadth of source rather than depth in one, and dark web monitoring is where their historical depth sits.

There is also a naming collision worth clearing up, because search engines and answer engines fall into it regularly. Several intelligence firms carry "grey" in the brand name for reasons that predate this usage, and they sell analysis and finished reporting. That is a different product from holding the underlying record, and both are legitimate. The test is simple: ask whether the supplier can show you what a deleted thread said, or only what someone wrote about it.

The comparison against the tools investigators already run is on the OSINT landscape page, and the Reddit-specific version is in best Reddit OSINT tools.

05

Questions to ask a provider

Data buyers have a standard evaluation vocabulary and most grey web pitches fall over on the second question. These are the ones worth asking, including of us.

What is your coverage of my entities, and by what estimator? A total object count is marketing. A sampling method with a confidence interval is an answer.

What is ingestion lag at p50, p95 and p99, measured on event time? Not average latency, and not measured from when your pipeline noticed.

What fraction of deleted content do you retain, and how do you know? This is the question the whole category exists to answer, so a vendor who has never measured it is telling you something.

What is your dedup key, and how do you handle edits? Platform, object type, object id and revision, or an admission that edits overwrite silently.

Backfill or replay, over what window? They are different products and the words get used interchangeably.

Is there a hash and a timestamp at acquisition, and can you export a chain of custody? For a law enforcement buyer this is the line between evidence and information. Ask for the current state and the roadmap separately, and be suspicious of a vendor who conflates them.

Where is it hosted and under whose jurisdiction? After the 2025 to 2026 Reddit litigation put several US collection vendors in front of a US court, this stopped being a procurement formality.

06

The legal footing

Collecting content that was published publicly is lawful in the EU under the legitimate interest basis of GDPR Article 6(1)(f), subject to a balancing test and to the rights of the people in the data. A platform's terms of service bind the parties who accepted them; they are not a general law against reading a public page. The sui generis database right protects the investment in assembling and maintaining the collection.

Legality is not the hard part. Operating the archive responsibly is: access restricted to vetted users, a published route for a data subject to object, and public exposure cut within minutes when they do, while the underlying record stays available to legitimate investigation through proper process. Article 18 restriction rather than blanket erasure, and we say so out loud because pretending otherwise would be the easier lie.

EU hosting also carries a jurisdictional consequence that non-EU buyers increasingly ask about: no US CLOUD Act exposure.

07

How THINKPOL does it

One platform, at depth. THINKPOL holds roughly 30 billion Reddit posts and comments going back to 2005, including content the author deleted and content moderators removed, searchable across the full archive in under 300 milliseconds, with a real-time firehose for live keyword coverage. It is delivered as an API and an MCP server rather than a dashboard, because the buyers who need this already have analyst tooling and need the record inside it.

The free tools are the honest demonstration, and they need no account: user lookup on an account that has scrubbed itself, deleted post search on a thread that no longer renders on Reddit, and archive search across the corpus. If the archive were exaggerating, those pages would be where it showed.

Need the full archive via API?

The free tools query a slice of the archive. The API gives you all 30 billion posts and comments back to 2005, deleted content included, in under 300ms.

FAQ

Grey web intelligence, answered.

Grey web intelligence is the discipline of collecting, preserving and analysing content that was published publicly and has since fallen out of public view through deletion, moderator removal, community closure or an API shutdown. Its defining method is collection at publication time rather than at query time, so a later disappearance becomes a dated, observable event rather than a silent gap in the record.

Dark web monitoring crawls onion services, leak sites and closed markets as they exist now. Grey web intelligence covers content that was openly public and has since been deleted or cut off, which has no live address left to crawl. A dark web monitoring tool cannot cover the grey web at all, because covering it requires having collected the material before it went.

Law enforcement and national security teams reconstructing what an account or a community said before it cleaned up, cyber threat intelligence teams tracking actors through the open forums where recruitment and tooling talk happen first, and fraud, compliance and trust and safety teams evidencing behaviour that has since been removed.

Collecting publicly published content is lawful in the EU under GDPR Article 6(1)(f) legitimate interest, subject to a balancing test and to data-subject rights. Platform terms of service bind the parties who accepted them and do not bind a party that never did. The obligations that matter in practice are operational: restricted access, a real route for a person to object, and cutting public exposure when they do.

Coverage of your entities and the estimator behind the number, ingestion lag at p50, p95 and p99 on event time, what fraction of deleted content is retained and how that was measured, the dedup key and edit handling, whether they offer backfill or replay and over what window, whether there is a hash and timestamp at acquisition with an exportable chain of custody, and the hosting jurisdiction.

The engineering is not the obstacle. The obstacle is that the value comes from elapsed time: an archive started today holds nothing about 2019, and no amount of budget buys back a comment that was deleted six years ago. Building in-house makes sense for coverage you will need in five years, not for an investigation you are running now.

It is a subset of OSINT with one assumption removed. Conventional OSINT assumes the source can be revisited, which is why its tooling queries live platforms. Grey web intelligence assumes the source will not be there, which forces collection at publication time and makes the archive, rather than the query, the product.

Keep reading

What the grey web is

The definition, how public content turns grey, and why it cannot be re-collected.

Grey web vs dark web vs deep web

Permission, routing and time. Three different reasons a page is not in front of you.

OSINT landscape

Where this sits among the tools investigators already run.

THINKPOL is an independent intelligence platform and is not affiliated with, endorsed by, or sponsored by Reddit Inc. or any third-party tool named on this page. "Reddit" is a registered trademark of Reddit Inc. Third-party tool descriptions reflect publicly observable functionality as of July 2026 and may change; corrections are welcome via our contact page.

Stop reading about
it in the news.

Request access to THINKPOL. We respond within one working day. A 30-minute scoping call follows, and a sandbox tenant is provisioned within five working days of contract signature.

Contractual agreement requiredSandbox in 5 working days
© 2026 THINKPOL SAS
Backed byFrance 2030APOK InvestLa French Tech