DEV Community
Follow
24 of the 30 duplicate groups were not duplicates
A live audit of AI-agent assets identified 30 groups of potential duplicates based solely on title matching. Content hashing, a more robust method, successfully resolved 24 of these groups. Five groups were confirmed as genuine duplicates, meaning identical content existed twice in the database. One group, however, contained rows with no content at all, making them unjudgeable.The audit evaluated duplicates based on content hashing, body similarity, provenance, and creation time. Beyond true duplicates, six groups had the same title but represented entirely different assets. Seven groups comprised distinct verification probes, synthetic assets used for testing pipelines. This left four groups, totaling 12 rows, with empty bodies, indicating a write-path incident rather than a duplication issue.Furthermore, the provenance field, intended to track asset origin, was found to be unreliable. Three pairs of assets had identical provenance references but completely unrelated content. This highlights the danger of relying on a single data point for deduplication. Ultimately, 80% of titles that appeared duplicated were not, and conversely, assets with different titles could be identical. The audit emphasizes the importance of using multiple signals, like content hash and provenance verification, for accurate asset identity judgement, rather than relying on title alone.