# What the grey web is The **grey web** (US spelling: **gray web**) is the layer of the public web that is technically open but practically unreachable: content that was published in the clear, then deleted by its author, removed by a moderator, locked behind a rate limit, or cut off when a platform closed its API. It needs no special software, which separates it from the dark web, and no credentials, which separates it from the deep web. What defines it is that it has fallen out of public view while remaining public in nature. It sits inside the surface web rather than beneath it: every page in it was indexed once. ## Where the term comes from The phrase predates this definition. Britannica describes the gray web as a part of the surface internet that needs no special service to reach but is associated with illegal activity. Bitsight, which ranks first for the phrase, narrows it to the part of the surface web used by fraudsters: forums for cracking tools and card fraud, open because nobody bothers to hide. That is accurate about a place and incomplete as a category, since a fraud forum is simply indexed surface web. What needs a name is the condition those pages share with a great deal of ordinary content: it does not stay. The fraud forums are one neighbourhood; the decay of the record is the category, and that is the definition used here. ## Three different blockers, not three layers - **Surface web**: nothing blocks it. Indexed, open in any browser. - **Deep web**: blocked by permission (login, paywall, query form). Credentials retrieve it. - **Dark web**: blocked by routing (Tor or equivalent). The right client reaches it. - **Grey web**: blocked by time. It was public and indexed, then it went. Nothing retrieves it; only a capture made before deletion survives. ## Grey web data **Grey web data** is the record of grey web content: posts, comments, profiles and community metadata captured while publicly visible and retained after they stopped being retrievable at source. Its defining property is that it cannot be re-collected. The constraint is not bandwidth, it is when collection started. ## Grey web intelligence **Grey web intelligence** is the discipline of collecting, preserving and analysing grey web data for investigative use. It differs from conventional OSINT because it does not assume the source will still be there: collection happens at publication time rather than at query time, so a deletion becomes an observable dated event rather than a silent gap in the record. ## Examples A Reddit comment its author deleted an hour after posting. A thread a moderator removed. A subreddit that went private or was banned, taking every discussion inside it out of view at once. A forum that shut down and took its address with it. A suspended account, whose history disappears from every live tool at once. A carding forum that was openly indexed until it was taken down. Each was readable by anyone at the time; none is reachable by anyone now, at any price, unless a copy was made while it was up. ## How content turns grey Author deletion, moderator removal, a community going private or banned, an API closing, or a platform shutting down. On a sample of Reddit comments recaptured seven days after collection, roughly one in twenty had already left public view, most removed by moderation rather than by their author. First measurement, small sample, still being consolidated. ## Legality Collecting publicly published content is lawful in the EU under GDPR Article 6(1)(f) legitimate interest, subject to a balancing test and to data-subject rights. Platform terms of service do not bind a party that never accepted them. The hard part is operating the archive responsibly: restricted access, honouring objections by cutting public exposure. THINKPOL operates a grey web archive of 30 billion Reddit posts and comments back to 2005, including author-deleted and moderator-removed content, with sub-300ms search over a REST API and MCP. Source: https://think-pol.com/grey-web