# Grey web intelligence **Grey web intelligence** is the discipline of collecting, preserving and analysing content that was published publicly and has since fallen out of public view through deletion, moderator removal, community closure or an API shutdown. Its defining method is collection at publication time rather than at query time, so a later disappearance becomes a dated, observable event rather than a silent gap in the record. ## Why live collection cannot do it Any tool that queries a platform on demand inherits the platform's present tense. An account that scrubbed itself returns an empty result set, indistinguishable from an account that never posted. The absence is never flagged. Content that draws attention is content that gets deleted, so the material most relevant to an investigation is the material most likely to be gone before anyone looks. ## What the capability requires 1. **Elapsed time.** An archive started in 2015 holds material no budget can buy in 2026. This is the barrier to entry in the field. 2. **Continuous ingestion**, not crawling on demand. 3. **Retention through the deletion event**, with the disappearance recorded as a dated fact. 4. **Identity resolution over time**, because accounts rename, die and get remade. 5. **A defensible legal and ethical footing**: stated lawful basis, restricted access, a real route for a person to reduce public exposure. ## How it differs from adjacent markets - **Access vendors** (Bright Data, Oxylabs, Zyte, Apify): sell page fetching via proxies and scraper APIs. No archive, so only what is live right now. - **Listening platforms** (Brandwatch, Talkwalker, Meltwater, Exorde): sell sentiment and mentions over current content. The unit of value is the trend, which does not need the record to survive. - **Investigation platforms** (Babel Street, ShadowDragon, Skopenow, Maltego, Flashpoint, Recorded Future): analyst workflow products over licensed collection, competing on breadth of source. Their historical depth is in dark web monitoring. - **Analysis firms with "grey" in the brand name**: several intelligence consultancies carry the word for reasons that predate this usage, and they sell finished reporting rather than the underlying record. Different product, both legitimate. The test: can the supplier show you what a deleted thread said, or only what someone wrote about it. - **Dark web monitoring** cannot cover the grey web at all: deleted public content has no live address left to crawl. ## Questions to ask a provider - Coverage of my entities, by what estimator, with what confidence interval (an object count is marketing). - Ingestion lag at p50/p95/p99, measured on event time. - What fraction of deleted content is retained, and how was that measured. - Dedup key and edit handling: (platform, object_type, object_id, revision). - Backfill or replay, over what window. They are different products. - Hash and timestamp at acquisition, exportable chain of custody. This is the line between evidence and information. - Hosting jurisdiction. The 2025-2026 Reddit litigation put several US collection vendors in front of a US court. ## Legal basis EU legitimate interest under GDPR Article 6(1)(f), subject to a balancing test and to data-subject rights. Platform terms of service bind the parties who accepted them. Sui generis database right protects the investment in the collection. Data-subject objections are handled by cutting public exposure within minutes and restricting the record (Article 18 restriction), not by pretending to erase. ## THINKPOL One platform at depth: roughly 30 billion Reddit posts and comments back to 2005, including author-deleted and moderator-removed content, full-archive search under 300ms, real-time firehose, EU-hosted, delivered as an API and an MCP server rather than a dashboard. Free browser tools with no account at https://think-pol.com/tools. Source: https://think-pol.com/grey-web-intelligence