Product

Reddit data API: the archive, not a proxy.

Most Reddit APIs read reddit.com on your behalf, which caps them at roughly a thousand items per listing and leaves them blind to anything deleted. THINKPOL serves a private archive of 30 billion posts and comments back to 2005, collected at publication time.

Short answer

If you need what Reddit currently shows, a live proxy API will do. If you need what was posted and then removed, you need an archive, because a proxy cannot return content that is no longer on the page it reads. THINKPOL covers 30 billion posts and comments back to 2005, retains author-deleted and moderator-removed content, searches the full corpus in under 300 milliseconds, and makes new content queryable in roughly ten seconds. Access is a documented REST API, also exposed over MCP, licensed annually. See pricing and licensing or request access.
01

What a Reddit data API has to answer

Teams buying Reddit data are usually trying to answer one of two questions, and they are not the same question. The first is what is being said right now: which accounts are posting about a brand, a vulnerability, a fraud technique or a person, and what has changed in the last few minutes. The second is what was said and then taken down: the comment an author deleted an hour after posting it, the thread a moderator removed, the account that scrubbed its history before anyone thought to look.

Most products on the market answer only the first. They read reddit.com on your behalf, normalise the response and hand it back as JSON. That is a legitimate product, and if your question is about live conversation it may be all you need. It cannot answer the second question at all, because the thing you are asking about is no longer on the page it reads.

Reddit's own Data API is the reference point here. It is quoted publicly at $12,000 per year plus $0.24 per 1,000 calls, and it serves what reddit.com currently shows. You are paying for authorised access to the present tense.

02

Proxy or archive: the thousand-item ceiling

This is the structural point, and it decides which category of product you actually need. Reddit caps any single public listing or feed at roughly 1,000 items. Ask for a subreddit's history, a user's post history or a search result set, and you get the most recent thousand entries and then the pagination stops. That ceiling is imposed at the source.

Anything that queries Reddit live inherits it. It does not matter how good the scraping layer is, how many residential proxies sit behind it or how clean the schema is: a proxy cannot return the 1,001st item, because the endpoint it depends on will not serve it. The honest vendors in that category say so in their own documentation. Some state plainly that they are not an archive and hold no database of old posts.

The second inherited limit is deletion. A live proxy returns what exists at the moment of the request. Content deleted by its author or removed by a moderator is gone from the response, and no amount of retry logic brings it back. For a marketing dashboard that is an inconvenience. For an investigation it is the case, because the material that gets deleted is disproportionately the material worth reading.

An archive is a different architecture, not a better proxy. It collects at publication time and keeps what it collected, so the ceiling and the deletion problem both disappear: the record exists in a database you control, independent of what reddit.com is willing to serve today. When you are comparing vendors, that single question separates them more reliably than any feature list.

03

What is in the archive

THINKPOL is a private archive of 30 billion Reddit posts and comments going back to 2005, collected continuously and still collecting. It was not rebuilt from public dumps, so it does not carry the gaps those dumps carry.

Because collection happens at publication time, content later deleted by its author or removed by moderators is retained. Full-text search across the whole corpus returns in under 300 milliseconds, which means the archive is usable interactively rather than as an overnight batch job. Identity resolution across accounts is part of the product, so an account is a starting point rather than the boundary of an enquiry.

You can see the corpus before you talk to anyone. The free archive search runs against the same data with no account, and how the free archive tools compare covers the non-commercial options honestly, including where they are the right answer.

04

The live firehose, and when you need it instead

The archive answers questions about the past. The firehose answers questions you have not asked yet. It is a separate product line: you register the keywords, accounts or terms you care about, and new matching content becomes queryable in roughly ten seconds.

The distinction matters for how you build. If your workflow is an analyst opening a case and reconstructing what happened, you are querying the archive. If your workflow is a detection pipeline that has to fire while something is still happening, a leak appearing, a credential-stuffing method circulating, a coordinated campaign starting, you need the firehose feeding it and the archive behind it for context on whoever shows up.

Most buyers end up taking both, because an alert without history is a notification rather than intelligence. They are priced and sized separately, so you can start with one.

05

Integration: REST, MCP and the shape of a first query

Access is a documented REST API. You authenticate, you issue a search across the corpus with the filters you would expect on a text index, and you get structured results back with the metadata needed to cite them later. There is no console to click through and no proprietary query language to learn before you can evaluate whether the data is any good.

The same surface is exposed over MCP, which matters if you are building agent tooling rather than a dashboard. An agent can query the archive directly as a tool, which removes the layer of glue code most teams write to make a REST API legible to a model.

In practice, an evaluation starts the same way every time. You take a case you already know the answer to, something your current source handled badly, and you run it against the archive to see whether the missing material is there. That is the test worth running, and it is the one we set up during a trial rather than asking you to take coverage claims on trust.

06

What it costs, and who we sell to

The two product lines are metered differently, because they are different things. Archive access is measured against query volume, with the unit rate falling as committed volume rises. Firehose access is measured against the number of terms you monitor and how long you monitor them. Both sit inside an annual agreement rather than a self-service signup, so the figures depend on scope. The current structure is on the pricing and licensing page.

Access is contractual and it is gated. We ask what the data will be used for, and we say no to some enquiries. That is a deliberate constraint of selling into law enforcement and government: those buyers cannot work with a supplier that sells to anyone. THINKPOL is built in France and hosted in the EU, with no US CLOUD Act exposure, which is a procurement requirement for European public-sector work rather than a marketing position.

The people who buy it are law enforcement, national security, cyber threat intelligence, fraud and compliance, trust and safety, and academic research. The common thread is that they need the record that Reddit no longer shows, which is the discipline we describe as grey web intelligence. If that is your problem, you can request access and we will scope it against a real case.

Need the full archive via API?

The free tools query a slice of the archive. The API gives you all 30 billion posts and comments back to 2005, deleted content included, in under 300ms.

FAQ

The API, the data, the contract.

It depends on whether you need the present or the past. Reddit's own Data API and the live proxy APIs built on it serve what reddit.com currently shows, subject to a ceiling of roughly 1,000 items per listing, and they cannot return deleted or removed content. THINKPOL is an archive rather than a proxy: 30 billion posts and comments back to 2005, collected at publication time, so deleted and removed material is retained and there is no listing ceiling. If your questions are about live conversation only, a proxy may be sufficient. If they are about what was said and then taken down, an archive is the only category that answers them.

Yes. THINKPOL exposes a private archive of 30 billion Reddit posts and comments going back to 2005 through a documented REST API, with full-text search across the whole corpus returning in under 300 milliseconds. The same surface is available over MCP for agent tooling. Reddit's own API is not a route to historical data at that depth, because it serves what the site currently shows.

THINKPOL is licensed rather than sold self-service. Archive access is metered against query volume, with the unit rate falling as committed volume rises, and the live firehose is metered against the number of terms you monitor and for how long. Both sit inside an annual agreement, so the number depends on scope. The current structure is set out on the pricing page, and a scoping call is the fastest way to get a figure for your use case.

Yes, within the archive. Collection happens at publication time, so content later deleted by its author or removed by a moderator is retained and remains queryable. Any product that reads reddit.com live at request time cannot do this, because the content is no longer on the page it reads.

A firehose is a live stream of new Reddit content matched against terms you register in advance, with new matching content queryable in roughly ten seconds. You need it if your workflow has to react while something is still happening, such as a leak appearing or a campaign starting. If you are reconstructing events after the fact, the historical archive is the right product.

THINKPOL is built in France and hosted in the EU, with no US CLOUD Act exposure. Access is contractual and vetted rather than self-service: buyers are law enforcement, national security, cyber threat intelligence, fraud and compliance, trust and safety teams, and academic researchers. We ask what the data will be used for before granting access.

Keep reading

Pricing and licensing

How archive access and firehose monitoring are licensed, and what a scoping call covers.

Grey web intelligence

The discipline behind the data: why the public record that disappears is the part worth keeping.

Best Reddit OSINT tools

Every category of Reddit tooling compared, including the free ones we do not charge for.

THINKPOL is an independent intelligence platform and is not affiliated with, endorsed by, or sponsored by Reddit Inc. or any third-party tool named on this page. "Reddit" is a registered trademark of Reddit Inc. Third-party tool descriptions reflect publicly observable functionality as of July 2026 and may change; corrections are welcome via our contact page.

Stop reading about
it in the news.

Request access to THINKPOL. We respond within one working day. A 30-minute scoping call follows, and a sandbox tenant is provisioned within five working days of contract signature.

Contractual agreement requiredSandbox in 5 working days
© 2026 THINKPOL SAS
Backed byFrance 2030APOK InvestLa French Tech