Reddit Deletion Statistics 2026: The THINKPOL Deletion Index
THINKPOL's Reddit Deletion Index: 2.9% of captured comments are gone within a month, 95% CI 1.9% to 4.6%, n = 616. Method, cohorts and limits.

THINKPOL runs a live archive of Reddit. The Deletion Index is our measurement of how much of what we capture stops being publicly retrievable on Reddit afterwards.
About 2.9 percent of the human comments THINKPOL captures are no longer publicly retrievable on Reddit within a month of capture. The 95 percent confidence interval is 1.9 to 4.6 percent, and the figure is pooled from n = 616 scored comments. Do not quote the percentage without the interval and the n.
This page gives that pooled figure, the three cohorts it pools, the method behind them and the limits of the measurement. It does not present a trend over time, because this measurement cannot yet support one.
Last verified: August 2026.
What does the Reddit Deletion Index measure?
The Deletion Index measures what share of the Reddit comments THINKPOL captured on a given day can no longer be retrieved publicly on Reddit a fixed number of days later.
It works by cohort. We take what the archive captured inside one archive window, wait a set number of days, then check each of those comments again on live Reddit. The index is the share that has become unreachable.
Bot comments are tallied separately and excluded from every figure on this page, because an AutoModerator post is never deleted by its author. Every percentage here is human-authored comments only.
How is the Deletion Index measured?
Each cohort draws 300 comments from a single archive window and rechecks every one of them at its Reddit permalink.
1. Select the window. Pull the comments our archive captured during a 30 minute window exactly N days before the run, using the archive search term "the" in word mode.
2. Spread the sample. Sample across the whole pulled range rather than taking the first page, so the draw is not dominated by a single minute of traffic.
3. Recheck on live Reddit. Fetch each comment at its old.reddit.com permalink from one IP, with a single global pacer and adaptive backoff on HTTP 429. The fetcher carries an over18 cookie, without which every NSFW comment reads as gone.
4. Classify each result. alive, removed (the body reads [removed]), deleted (the body reads [deleted]), missing (Reddit does not know the comment id and the permalink redirects to the thread root), post_gone (the parent submission is unreachable), age_gated, unusable, fetch_error.
5. Score. The denominator excludes fetch errors, age-gated pages and unusable pages. Anything we cannot confidently classify as a disappearance is counted as still present, which biases the index downward rather than upward.
6. Separate the bots. AutoModerator and similar accounts are tallied apart from the human figure and reported separately.
The run parameters are the cohort age and the sample size, and nothing else. Anyone holding an equivalent archive can reproduce the design. You can query our archive directly with the THINKPOL Reddit search tool.
What is the headline figure?
About 2.9 percent of human comments THINKPOL captures are no longer publicly retrievable on Reddit within a month of capture, 95 percent confidence interval 1.9 to 4.6 percent, n = 616 scored comments.
That figure pools the three cohorts measured on 18 August 2026: 18 disappearances out of 616 scored human comments, which is 2.92 percent, with a Wilson 95 percent interval of 1.86 to 4.57 percent.
Pooling is the only defensible way to read this run. Each cohort on its own turns on five to seven events, which is too few to support a number anyone should act on. Pooled, the three draws give one rate for the first month of a comment's life with an interval narrow enough to be useful.
What did each cohort show?
The three cohorts scored 3.52, 2.88 and 2.39 percent, and their confidence intervals overlap almost completely.
Every row carries its n and its Wilson 95 percent interval. None of these percentages means anything without them.
1 day cohort. Archive window 17 August 2026, 20:23 to 20:53 UTC. n = 199 scored human comments, 7 gone. 3.52 percent, Wilson 95 percent interval 1.71 to 7.08 percent.
7 day cohort. Archive window 11 August 2026, 19:42 to 20:12 UTC. n = 208 scored human comments, 6 gone. 2.88 percent, Wilson 95 percent interval 1.33 to 6.15 percent.
30 day cohort. Archive window 19 July 2026, 21:11 to 21:41 UTC. n = 209 scored human comments, 5 gone. 2.39 percent, Wilson 95 percent interval 1.03 to 5.48 percent.
Pooled, all three cohorts. n = 616 scored human comments, 18 gone. 2.92 percent, Wilson 95 percent interval 1.86 to 4.57 percent. This is the headline figure.
Every cohort started from a draw of 300 comments. Bot accounts accounted for 23, 28 and 21 of those draws. Reddit rate-limited our fetcher on 83, 67 and 72 comments respectively, which is 22 to 28 percent of each draw, and those are excluded from the denominator. A comment we could not fetch is a measurement we do not have, not a disappearance.
The intervals are computed with the Wilson score method at z = 1.96, from the counts above and nothing else. The disappearances were not concentrated in any community: in each of the three cohorts, every affected subreddit appeared exactly once.
Can a trend be read across one, seven and thirty days?
No. The three cohort figures are independent draws of different content, not one cohort followed over time, so they cannot be read as a trend.
Deletion is cumulative. A comment that is gone at one day is still gone at thirty days, so a rate that falls with age is impossible for a single tracked cohort. The 1 day, 7 day and 30 day figures here come from three different sets of comments, captured on three different dates and rechecked once each. That is why they can move in any direction.
Their confidence intervals overlap almost completely: 1.71 to 7.08 percent, 1.33 to 6.15 percent, 1.03 to 5.48 percent. Each of those ranges contains all three point estimates. At this sample size the measurement cannot distinguish one day from one month.
The apparent decline with age is sampling noise, not a real effect, and we will not present it as one. With five to seven events per cohort, one comment either way moves a figure by roughly half a percentage point. Separating 3.5 percent from 2.4 percent at 95 percent confidence and 80 percent power needs roughly 3,700 scored comments per cohort. We have about 200.
So this page presents no survival curve, and any citation of these numbers as a decline over time is a misreading of them.
Who removes the content, the author or Reddit?
At these horizons the disappearance is dominated by Reddit-side suppression rather than by authors deleting their own comments.
The 7 day cohort is the clearest case. Of its 6 disappearances, 4 were comments whose parent post disappeared and 2 were comments Reddit no longer serves at their own permalink. None came back as author-deleted. None came back as moderator-removed-in-place. The author acted in zero of the six cases.
Gone means a comment we captured can no longer be read at its Reddit permalink by a reader who is not logged in. That covers four distinct outcomes, and they are not the same thing.
The parent post is gone. The comment may still exist somewhere in Reddit's database, but the thread that held it is unreachable, so the comment is unreachable with it. This was the largest category at 1 and 7 days: 6 of 7 disappearances at one day, 4 of 6 at seven days.
The comment is missing. Reddit does not recognise the comment id and the permalink redirects to the thread root. In practice this is mostly moderator removal, because a mod-removed comment is often absent from the logged-out tree entirely rather than shown as [removed]. Two cases at seven days, four at thirty days.
The author deleted it. The body reads [deleted]. One comment in the 1 day cohort and one in the 30 day cohort, out of 18 disappearances in total.
The moderators removed it in place. The body reads [removed]. No comment in any of the three cohorts landed in this category, which is consistent with removals surfacing as missing instead.
This is the part of the measurement that matters for an investigator. When a comment stops being retrievable it is usually because something on Reddit's side removed the path to it, not because the author changed their mind. Checking Reddit alone shows an absence with no explanation and no record that anything was ever posted there. An archive that holds the original public capture preserves the record of what was publicly posted and when, which is what makes the absence readable and what makes a timeline hold up under review.
Missing means suppressed from public view. It does not mean the author deleted anything, and we do not describe it that way. If you want the practical version of this, see our guide on how to view deleted Reddit posts.
Why can only an independent archive measure this?
Measuring what disappeared from Reddit requires holding the original capture, because Reddit does not publish a record of what it removes.
Query Reddit today and you see what survived. There is nothing in that answer to compare against, so the deletion rate is invisible from Reddit's own surface. The measurement needs two things at once: a record of what was publicly visible at time T, and a recheck of the same items at time T plus N days.
That is what an independent archive is, and it is why this index is a by-product of running the collection rather than a research project bolted onto it. The mechanics of that capture are covered in our explainer on Reddit archive search.
Everything measured here was publicly visible on Reddit at the moment we captured it. Nothing on this page comes from private, paywalled or unlawfully obtained material.
What are the limitations of this measurement?
Six limitations apply, and the first of them biases the published figure low.
Rate limiting removes a quarter of each draw. Reddit rate-limits our collector. Between 22 and 28 percent of each 300 comment draw failed to fetch: 83, 67 and 72 comments. Those failures are excluded from the denominator. If any of them are in fact disappearances, the published figure is biased low.
Each cohort is one half hour of Reddit. Each cohort was drawn from a single 30 minute slice of archive time, so it is clustered on whatever was happening during that half hour. It is not a random sample of Reddit.
The sample skews English. Comments are drawn with the search term "the" in word mode, so the sample skews English-language and omits any comment that does not contain that word.
The sample is far too small to compare cohorts. Separating 3.5 percent from 2.4 percent at 95 percent confidence and 80 percent power needs roughly 3,700 scored comments per cohort, not 200. The three cohorts turn on 7, 6 and 5 events.
The NSFW interstitial. NSFW subreddits serve an over-18 interstitial to a logged-out fetcher. Without an over18 cookie, every NSFW comment reads as gone. This was found and fixed before these runs. It is the single largest way a measurement of this kind can be got wrong, and any third-party deletion statistic that does not mention it should be treated with suspicion.
Residential proxies are not a workaround. Routing the recheck through a residential proxy returned 6 of 6 comments as missing during testing. That was a login wall, not deletion. The tool now marks any page containing neither a submission node nor a comment node as unusable and drops it from the denominator instead of counting it as a disappearance.
What are we doing about the sample size?
A pooled 7 day measurement now runs nightly, and each night draws a different 6 hour slice of archive time.
Pooling nightly runs addresses both the sample size and the window clustering at once, because different nights draw different slices and the pool stops being one half hour of Reddit. We will update this page when the pool passes 3,000 scored comments. We are not putting a date on that.
Until then this page is a first reading, not a series. Each update will be published at this URL, including any run that contradicts this one. The URL carries no year, so it is updated in place rather than replaced.
How should this page be cited?
Cite the Reddit Deletion Index by name, year and URL, and quote the confidence interval and the n alongside any percentage taken from it.
Cite this
Suggested citation: THINKPOL, Reddit Deletion Index, first readings, August 2026, https://think-pol.com/blogs/reddit-deletion-statistics
Page URL: https://think-pol.com/blogs/reddit-deletion-statistics
Measurement date: 18 August 2026. Cohort ages: 1, 7 and 30 days. Method: cohort recheck of archived comments at their Reddit permalinks, Wilson 95 percent intervals.
If you quote one number, quote this one: about 2.9 percent of the human comments THINKPOL captures are no longer publicly retrievable on Reddit within a month of capture, 95 percent confidence interval 1.9 to 4.6 percent, n = 616 scored comments.
See what you've been missing.
Real-time grey web intelligence before threats materialise. Vetted access only.



