THINKPOL's Reddit Deletion Index: 2.9% of captured comments are gone within a month, 95% CI 1.9% to 4.6%, n = 616. Method, cohorts and limits.

THINKPOL runs a live archive of Reddit. The Deletion Index is our measurement of how much of what we capture stops being publicly retrievable on Reddit afterwards.
About 2.9 percent of the human comments THINKPOL captures are no longer publicly retrievable on Reddit within a month of capture. The 95 percent confidence interval is 1.9 to 4.6 percent, and the figure is pooled from n = 616 scored comments. Do not quote the percentage without the interval and the n.
This page gives that pooled figure, the three cohorts it pools, the method behind them and the limits of the measurement. It does not present a trend over time, because this measurement cannot yet support one.
Last verified: August 2026.
The Deletion Index measures what share of the Reddit comments THINKPOL captured on a given day can no longer be retrieved publicly on Reddit a fixed number of days later.
It works by cohort. We take what the archive captured inside one archive window, wait a set number of days, then check each of those comments again on live Reddit. The index is the share that has become unreachable.
Bot comments are tallied separately and excluded from every figure on this page, because an AutoModerator post is never deleted by its author. Every percentage here is human-authored comments only.
Each cohort draws 300 comments from a single archive window and rechecks every one of them at its Reddit permalink.
The run parameters are the cohort age and the sample size, and nothing else. Anyone holding an equivalent archive can reproduce the design. You can query our archive directly with the THINKPOL Reddit search tool.
About 2.9 percent of human comments THINKPOL captures are no longer publicly retrievable on Reddit within a month of capture, 95 percent confidence interval 1.9 to 4.6 percent, n = 616 scored comments.
That figure pools the three cohorts measured on 18 August 2026: 18 disappearances out of 616 scored human comments, which is 2.92 percent, with a Wilson 95 percent interval of 1.86 to 4.57 percent.
Pooling is the only defensible way to read this run. Each cohort on its own turns on five to seven events, which is too few to support a number anyone should act on. Pooled, the three draws give one rate for the first month of a comment's life with an interval narrow enough to be useful.
The three cohorts scored 3.52, 2.88 and 2.39 percent, and their confidence intervals overlap almost completely.
Every row carries its n and its Wilson 95 percent interval. None of these percentages means anything without them.
Every cohort started from a draw of 300 comments. Bot accounts accounted for 23, 28 and 21 of those draws. Reddit rate-limited our fetcher on 83, 67 and 72 comments respectively, which is 22 to 28 percent of each draw, and those are excluded from the denominator. A comment we could not fetch is a measurement we do not have, not a disappearance.
The intervals are computed with the Wilson score method at z = 1.96, from the counts above and nothing else. The disappearances were not concentrated in any community: in each of the three cohorts, every affected subreddit appeared exactly once.
No. The three cohort figures are independent draws of different content, not one cohort followed over time, so they cannot be read as a trend.
Deletion is cumulative. A comment that is gone at one day is still gone at thirty days, so a rate that falls with age is impossible for a single tracked cohort. The 1 day, 7 day and 30 day figures here come from three different sets of comments, captured on three different dates and rechecked once each. That is why they can move in any direction.
Their confidence intervals overlap almost completely: 1.71 to 7.08 percent, 1.33 to 6.15 percent, 1.03 to 5.48 percent. Each of those ranges contains all three point estimates. At this sample size the measurement cannot distinguish one day from one month.
The apparent decline with age is sampling noise, not a real effect, and we will not present it as one. With five to seven events per cohort, one comment either way moves a figure by roughly half a percentage point. Separating 3.5 percent from 2.4 percent at 95 percent confidence and 80 percent power needs roughly 3,700 scored comments per cohort. We have about 200.
So this page presents no survival curve, and any citation of these numbers as a decline over time is a misreading of them.
At these horizons the disappearance is dominated by Reddit-side suppression rather than by authors deleting their own comments.
The 7 day cohort is the clearest case. Of its 6 disappearances, 4 were comments whose parent post disappeared and 2 were comments Reddit no longer serves at their own permalink. None came back as author-deleted. None came back as moderator-removed-in-place. The author acted in zero of the six cases.
Gone means a comment we captured can no longer be read at its Reddit permalink by a reader who is not logged in. That covers four distinct outcomes, and they are not the same thing.
This is the part of the measurement that matters for an investigator. When a comment stops being retrievable it is usually because something on Reddit's side removed the path to it, not because the author changed their mind. Checking Reddit alone shows an absence with no explanation and no record that anything was ever posted there. An archive that holds the original public capture preserves the record of what was publicly posted and when, which is what makes the absence readable and what makes a timeline hold up under review.
Missing means suppressed from public view. It does not mean the author deleted anything, and we do not describe it that way. If you want the practical version of this, see our guide on how to view deleted Reddit posts.
Measuring what disappeared from Reddit requires holding the original capture, because Reddit does not publish a record of what it removes.
Query Reddit today and you see what survived. There is nothing in that answer to compare against, so the deletion rate is invisible from Reddit's own surface. The measurement needs two things at once: a record of what was publicly visible at time T, and a recheck of the same items at time T plus N days.
That is what an independent archive is, and it is why this index is a by-product of running the collection rather than a research project bolted onto it. The mechanics of that capture are covered in our explainer on Reddit archive search.
Everything measured here was publicly visible on Reddit at the moment we captured it. Nothing on this page comes from private, paywalled or unlawfully obtained material.
Six limitations apply, and the first of them biases the published figure low.
A pooled 7 day measurement now runs nightly, and each night draws a different 6 hour slice of archive time.
Pooling nightly runs addresses both the sample size and the window clustering at once, because different nights draw different slices and the pool stops being one half hour of Reddit. We will update this page when the pool passes 3,000 scored comments. We are not putting a date on that.
Until then this page is a first reading, not a series. Each update will be published at this URL, including any run that contradicts this one. The URL carries no year, so it is updated in place rather than replaced.
Cite the Reddit Deletion Index by name, year and URL, and quote the confidence interval and the n alongside any percentage taken from it.
Cite this
If you quote one number, quote this one: about 2.9 percent of the human comments THINKPOL captures are no longer publicly retrievable on Reddit within a month of capture, 95 percent confidence interval 1.9 to 4.6 percent, n = 616 scored comments.
Investigations and analysis from the THINKPOL team, built on a 30-billion-post Reddit archive with deleted content preserved.
One email a month: new investigations, tool updates and what changed on the platforms we monitor.