All updates
Analytics

Which files are duplicated, not just how much

The Savings Finder told you how many terabytes were duplicated. Now it shows you the files — grouped, described in plain language, and linked straight to the object in your bucket.

"1.5 million objects are byte-identical copies" is a number you cannot act on. See the biggest on the duplicates recommendation opens the findings themselves: the largest groups, worst waste first, each one linked to the file in your browser.

Matched on content, not on name — and that turns out to be the point

Two objects count as duplicates when their checksum and size match. The name is not part of it, and in practice most duplicated bytes are files stored under names that do not match — the kind of thing you would never find by searching for a repeated filename.

Each group is described rather than just counted

Where the copies form a recognisable pattern, the finding says so:

  • *One file, filed once into each of 48 folders — Page 21 to Page 68* — a document covering a page range, written once under every page it contains.
  • *The same file in 122 folders, in 3 buckets* — a file that has spread across buckets, which usually means a copy nobody remembers making.
  • *184 different names across 89 folders* — identical content, unrelated names.

No delete button, and that is deliberate

A duplicate is not automatically waste. Something in your systems may be resolving a path to each of those copies, and collapsing them is a change at your end, not a cleanup we can do from a dialog. The findings tell you where to look; what to remove stays your decision.

Worth knowing before you act: for files uploaded in parts, the checksum covers the parts rather than the whole object, so it depends on the part size used at upload as well as the content. Matching checksum and size is strong evidence — confirm before deleting anything irreplaceable. Where a group is very large the dialog shows a sample of the copies and says so; the count itself is exact.