I recently decided to clean up years of files spread across old computers, external drives, backups and different stages of my life.
At first, I thought this was going to be a file-management project.
Create a better folder structure. Find the duplicates. Move everything into the right place. Delete what I no longer needed.
Simple.
Then I started looking at the actual data.
The original problem was bigger than the final archive makes it look. Across the source material, I was dealing with roughly 1.5 TB of information. An initial round of consolidation and cleanup brought that down to around 700 GB. By the stage I am writing about here, the verified working inventory contained 36,661 files and roughly 646.91 GB of data.
Construction documents. Personal files. Photos and videos. Music. Old software installers. Downloads. Backups from previous computers. Files that had been copied from one drive to another years ago. Files with meaningful names mixed with files whose names told me almost nothing.
The deeper I got into it, the more I realized that organizing a large archive safely is not really a folder problem.
It is an information problem.
And if I was going to automate any part of it, I needed to be very careful about the difference between a computer making a reasonable guess and a computer actually knowing what something is.
The first rule was simple: don't move anything yet
My first instinct was to start building the new folder structure and moving things into it.
Instead, I decided to inspect everything first.
That turned out to be one of the most important decisions in the project.
PowerShell made it possible to inventory an entire drive without modifying anything:
Get-ChildItem 'X:\ArchiveSource' -File -Recurse |
Select-Object FullName, Length, Extension, LastWriteTime
That immediately gives you a much clearer picture of what you actually have.
From there I could start answering basic questions:
How many files are there?
How much space do they occupy?
What types of files are most common?
Where are the largest folders?
How much of the storage consists of photos and videos?
How many files have the same name?
Where are the obvious backup folders?
A simple summary can already be useful:
$files = Get-ChildItem 'X:\ArchiveSource' -File -Recurse
[pscustomobject]@{
Files = $files.Count
SizeGB = [math]::Round(
($files | Measure-Object Length -Sum).Sum / 1GB,
2
)
}
The important part was that this stage was read-only.
I wanted to understand the archive before trying to improve it.
That may sound overly cautious, but once you are dealing with tens of thousands of files accumulated over years, blindly reorganizing everything is an excellent way to create a cleaner-looking mess.
I wanted a simple structure, not a complicated taxonomy
One mistake I wanted to avoid was creating hundreds of folders trying to describe every possible type of information.
That sounds organized in theory.
In practice, it becomes another system you have to maintain.
Instead, I started with broad categories.
Things like personal documents, photos and video, construction and professional files, media, reference material, installers, old backups and one especially important category:
Inbox.
The inbox was for anything that could not be confidently classified.
That ended up becoming one of my favourite ideas from the project.
A file did not need to be forced somewhere just because the system had to produce an answer.
If the classification was uncertain, uncertainty itself was allowed to be the result.
That is much safer than pretending every decision has the same level of confidence.
A filename is not enough
The next problem was classification.
A file extension can tell you something.
A .jpg is probably an image.
An .mp4 is probably video.
An .exe is probably an installer or application.
But other files are much less obvious.
A PDF might be a construction drawing, a specification, a personal document, an invoice, a manual, a book or an old scanned record.
The filename may help, but even that can be misleading.
I had work documents whose surrounding folders were more meaningful than the filename itself.
That meant classification could not rely on one signal.
The path mattered. The extension mattered. The filename mattered. And in some cases, the file would eventually need deeper inspection.
That led me to a much better way of thinking about automation:
Automate the obvious cases and surface the questionable ones.
Do not try to eliminate human review.
Use automation to reduce the amount of human review that is necessary.
Duplicate files were more complicated than I expected
Duplicates seemed like an easy place to reclaim storage.
Then I started thinking about what a duplicate actually is.
Suppose I have Report.pdf in two different folders.
Those filenames match.
That does not mean the files are identical.
One may have been updated. One might contain completely different information.
Even matching file sizes are not proof.
A much stronger way to compare files is with a cryptographic hash.
I used SHA-256:
Get-FileHash 'X:\ArchiveSource\Report.pdf' -Algorithm SHA256
A SHA-256 hash acts like a fingerprint of the file's contents.
Two files with identical SHA-256 values provide extremely strong evidence that their actual bytes are the same.
That meant I could start grouping files by their contents rather than their names.
Get-ChildItem 'X:\ArchiveSource' -File -Recurse |
ForEach-Object {
$hash = Get-FileHash $_.FullName -Algorithm SHA256
[pscustomobject]@{
Path = $_.FullName
Size = $_.Length
SHA256 = $hash.Hash
}
} |
Group-Object SHA256
That works.
But it also creates another problem.
Two files can contain identical bytes while still representing different pieces of history.
Identical files can have different context
Imagine the same drawing exists inside two different construction project folders.
The bytes may be completely identical.
Technically, it is duplicate data.
But those two locations may tell me that the drawing was relevant to both projects.
If I keep one copy and permanently forget that the other ever existed, I may save some space but lose information.
That changed the way I approached deduplication.
The hash answers one question:
Are these files byte-for-byte identical?
It does not answer:
Are these files interchangeable in every meaningful way?
Those are different questions.
The same concept applies to personal files.
A photo may appear in an old phone backup and in a manually created family folder. The photo itself is the same. The locations tell two different stories about where it came from and how I used it.
So I started treating the original location as information worth preserving.
The old file path became metadata
Normally a path is just how we reach a file:
D:\Old Computer\Documents\Projects\Example.pdf
During a migration, though, that path tells you something.
Old Computer might identify where the file came from.
Projects tells you something about its purpose.
The parent folders may contain context that would disappear if the file were simply dropped into Documents\PDF.
So instead of thinking only in terms of:
old location → new location
I started thinking about recording:
original location
file identity
file size
hash
proposed category
proposed destination
decision
result
verification
That was the point where the project started feeling less like cleaning a hard drive and more like building a small information-management system.
Planning and execution should be separate
This became another major rule.
A program should be able to inspect the entire archive and tell me what it would do without actually doing it.
I wanted a planning stage that could inventory files, calculate hashes, identify duplicates, propose categories, identify possible filename collisions, recommend destination paths and flag uncertain files.
And then stop.
NO MOVES
NO RENAMES
NO DELETES
Only after examining the proposed plan would I allow anything to change.
That separation gives you something incredibly valuable:
the ability to discover that your logic is wrong before it touches your files.
A move should not really be one operation
Windows makes moving a file look simple.
Drag it from Folder A to Folder B.
Done.
For important archives, I wanted a safer definition.
COPY
↓
VERIFY
↓
RECORD SUCCESS
↓
ONLY THEN CONSIDER REMOVING THE SOURCE
Suppose I copy a file to its new location.
I can calculate both hashes:
$sourceHash = (Get-FileHash $source -Algorithm SHA256).Hash
$destHash = (Get-FileHash $destination -Algorithm SHA256).Hash
if ($sourceHash -eq $destHash) {
'VERIFIED'
}
Now I know the destination contains the same bytes as the source.
But even then, I do not necessarily want the original automatically deleted.
Deleting the source is another decision.
That distinction makes the whole process much more resilient.
If the destination drive disconnects, the original still exists.
If a script crashes halfway through, the original still exists.
If I later realize my folder structure needs to change, the original still exists.
Storage is cheap compared with losing something that cannot be recreated.
I also stopped trusting drive letters as identities
Another problem became obvious once external drives were involved.
Windows may call a drive E: today and F: tomorrow.
Drive letters are locations assigned by the operating system.
They are not the identity of the storage itself.
That becomes especially important if an archive might eventually move between computers or onto network storage.
Instead of permanently thinking, "My archive is E:", I started thinking, "This is my archive, and Windows currently happens to mount it here."
That sounds like a small distinction, but it makes automation much safer.
The software should confirm that it is dealing with the storage it expects before performing important operations.
It should not assume that anything mounted at a particular drive letter must be the right drive.
The computer should prove its work
At one stage, an audit appeared to report 12 missing files.
That is the type of message that gets your attention very quickly when you are reorganizing important data.
But instead of assuming the files were gone, I investigated the audit itself.
It turned out that the comparison had run into an issue related to how some long Windows paths were being handled.
After fixing the checking logic and running the verification again, the result was:
0 actual missing files.
The planning stage also reported:
0 destination collisions in that verified comparison.
That experience taught me another valuable rule:
A diagnostic result is evidence, not absolute truth.
If a script says twelve files are missing, verify that the script actually looked for those twelve files correctly.
The measuring tool can fail too.
In other words:
verify the verification.
Long-running jobs should save their progress
Hashing tens of thousands of files takes time.
So does indexing, copying and validating large archives.
At first it is tempting to let a script run for an hour and generate one beautiful final report.
There is a problem with that design.
If the process fails near the end, you may lose all of the work you just performed.
A better system periodically saves its intermediate results.
$results = 'X:\Working\hash-checkpoint.csv'
Get-ChildItem 'X:\ArchiveSource' -File -Recurse |
ForEach-Object {
$hash = Get-FileHash $_.FullName -Algorithm SHA256
[pscustomobject]@{
Path = $_.FullName
Size = $_.Length
SHA256 = $hash.Hash
} |
Export-Csv $results -Append -NoTypeInformation
}
A more advanced version would batch writes for performance, but the principle remains the same:
do not make the final line of a long script the first moment your progress becomes durable.
If something stops halfway through, I want to resume from a known checkpoint rather than start from zero.
The system needed to be comfortable saying "stop"
As the project became more automated, I started thinking about failure conditions differently.
Imagine an external drive disappears halfway through a copy.
The wrong response is:
Try something else and keep going.
The safe response is:
Stop.
Keep the source.
Record where the operation failed.
Do not mark the file complete.
Resume only once the storage is available and has been verified again.
The same applies to a filename collision.
If Project Report.pdf already exists at the destination but has a different hash, automatically overwriting it would be unacceptable.
That is not a duplicate.
That is a conflict.
A good system should flag it for review.
The computer does not always need another clever fallback.
Sometimes stopping is the intelligent behaviour.
My definition of "organized" changed
When I started this project, I imagined success as a perfectly clean drive.
Every file in exactly the right folder.
No duplicates.
No clutter.
Beautiful structure.
I don't think that is the real goal anymore.
A genuinely useful archive is one where I can answer questions like:
Where did this file originally come from?
Has this file changed?
Do I have another identical copy?
Why was this file placed here?
Was this copy verified?
Did anything get overwritten?
Which files still need review?
Can I recover if the process fails?
That matters much more to me than whether every folder looks beautiful in File Explorer.
The biggest lesson was about automation, not storage
The more I worked on the archive, the more I realized the project had become a useful lesson in automation in general.
Automation is easy when the computer is allowed to be wrong.
The difficult part is building automation for information you actually care about.
That requires different rules.
Measure before changing things.
Preserve original data.
Separate planning from execution.
Use strong verification instead of assumptions.
Do not treat filenames as identity.
Do not treat identical bytes as identical context.
Keep uncertain cases for review.
Checkpoint long-running jobs.
Make important operations reversible.
And when the system is not confident, allow it to stop.
Those ideas started with a pile of old files.
But they apply to almost any system that is going to make decisions on a person's behalf.
The most useful automation is not the one that does the most things by itself.
It is the one I can trust to know the difference between something it understands, something it can verify, and something it should leave alone until I look at it.
This project started as storage cleanup. It ended up changing how I think about automation, verification and the systems I want to build next.
More from the blog →