
Python Automation: Scripts That Save You Hours
You know the Python basics: variables, loops, functions, a bit of file handling. This guide takes the next step and turns that knowledge into scripts that do real chores for you. By the end you will be able to sort a messy folder, rename files in bulk, process a CSV, run the whole thing on a schedule, and call it as a command from anywhere on your machine.
Everything here uses the standard library. No packages to install, nothing to break on the next upgrade. If you need a refresher on syntax first, the Python basics tutorial covers the ground this post assumes.
Picking a task worth automating
Not every repetitive job is a good candidate. A script pays for itself when the task is frequent, boring, and has rules you can write down. It is a poor fit when every case needs a judgement call, because you will spend more time patching exceptions than you ever spent doing the work by hand.
Before writing code, do the task manually once and narrate the rules out loud. "Anything ending in .pdf goes in documents. Screenshots go in images. Anything I do not recognise stays where it is." That sentence is your specification. If you cannot say it in three sentences, the task is not ready to be automated yet.
Be honest about the payoff too. A script that saves ten minutes a week is worth an hour of work. A script that saves ten minutes a year is a hobby project, which is fine, as long as you know that is what it is.
Step 1: Walk a folder with pathlib
pathlib replaces the old string juggling with os.path. A Path object knows its own name, extension, parent and size, and it behaves the same on Windows, macOS and Linux.
from pathlib import Path
downloads = Path.home() / "Downloads"
for item in sorted(downloads.iterdir()):
if item.is_file():
size_kb = item.stat().st_size / 1024
print(f"{item.name:40} {size_kb:8.1f} KB")
The slash operator joins path segments, so you never have to think about whether to write a backslash or a forward slash. item.suffix gives you .pdf, item.stem gives you the name without the extension, and item.parent gives you the containing folder.
Use iterdir() for one level and rglob("*.pdf") when you want to recurse into subfolders. Wrapping the result in sorted() does two useful things: it gives you a predictable order, and it materialises the listing into a list before you start changing the folder underneath yourself.
Step 2: Sort files into folders by type
Here is the file-organiser. Note the dry_run flag, which is the single most valuable habit in this whole article.
from pathlib import Path
BUCKETS = {
"images": {".jpg", ".jpeg", ".png", ".gif", ".webp"},
"documents": {".pdf", ".docx", ".txt", ".md"},
"archives": {".zip", ".tar", ".gz"},
}
def bucket_for(path):
suffix = path.suffix.lower()
for name, suffixes in BUCKETS.items():
if suffix in suffixes:
return name
return "other"
def unique_target(target):
if not target.exists():
return target
counter = 2
while True:
candidate = target.with_name(f"{target.stem}-{counter}{target.suffix}")
if not candidate.exists():
return candidate
counter += 1
def organise(source, dry_run=True):
for item in sorted(source.iterdir()):
if not item.is_file():
continue
target_dir = source / bucket_for(item)
target = unique_target(target_dir / item.name)
if dry_run:
print(f"{item.name} -> {target_dir.name}/{target.name}")
continue
target_dir.mkdir(exist_ok=True)
item.replace(target)
Three details that matter. suffix.lower() catches .JPG as well as .jpg. unique_target stops a second invoice.pdf from silently destroying the first one, which is exactly what a bare replace() would do. And mkdir(exist_ok=True) means a second run does not crash on a folder that already exists.
One caveat: Path.replace() is a rename, so it only works within the same filesystem. Moving files to another drive or an external disk raises OSError with a cross-device message. Use shutil.move(item, target) when the destination might live elsewhere. It is slower, because it copies and deletes, but it works everywhere.
Step 3: Batch-rename without wrecking your files
Renaming is where automation goes wrong most spectacularly, because a bad pattern applied to four hundred files is four hundred mistakes. Always print the plan before you execute it.
import re
from pathlib import Path
folder = Path("invoices")
pattern = re.compile(r"^scan_(\d{4})(\d{2})(\d{2})_(\d+)\.pdf$")
for item in sorted(folder.glob("*.pdf")):
match = pattern.match(item.name)
if not match:
print(f"skipped: {item.name}")
continue
year, month, day, index = match.groups()
new_name = f"{year}-{month}-{day}-invoice-{int(index):03d}.pdf"
print(f"{item.name} -> {new_name}")
# item.rename(item.with_name(new_name))
Run it with the rename line commented out, read the output, and only then uncomment it. Files that do not match the pattern are reported rather than skipped in silence, so you notice when your assumption about the naming scheme was wrong.
Two traps worth knowing. Renaming onto an existing name overwrites it on Unix and raises an error on Windows, so run the result through unique_target if collisions are plausible. And if you are shifting numbers up (file 1 becomes file 2, and so on), process the list in reverse order, otherwise each rename lands on the name of a file you have not moved yet.
Step 4: Read and write CSV with the standard library
The csv module handles quoting, embedded commas and line breaks inside fields. Hand-rolled line.split(",") handles none of that, and it will bite you the first time a customer has a comma in their address.
import csv
from pathlib import Path
rows = []
with Path("orders.csv").open(newline="", encoding="utf-8") as handle:
for row in csv.DictReader(handle):
row["total"] = float(row["quantity"]) * float(row["unit_price"])
rows.append(row)
large = [row for row in rows if row["total"] > 100]
fields = ["order_id", "customer", "total"]
with Path("large-orders.csv").open("w", newline="", encoding="utf-8") as handle:
writer = csv.DictWriter(handle, fieldnames=fields, extrasaction="ignore")
writer.writeheader()
writer.writerows(large)
newline="" is not optional. Leave it out and every row in your output file gets a blank line after it on Windows. Set encoding="utf-8" explicitly as well, because the default encoding depends on the machine, and a script that works on your laptop should not corrupt accented names on a colleague's.
extrasaction="ignore" lets you write a subset of the columns you read. Without it, DictWriter raises an error on any key that is not in fieldnames.
Step 5: Run it on a schedule
A script you have to remember to run is not automation. Both major operating systems ship a scheduler, and the concepts are the same on each: pick a program, pick arguments, pick a time.
On Linux and macOS, cron reads a table of lines where the first five fields are minute, hour, day of month, month and day of week. Edit it with crontab -e.
0 7 * * 1 /home/you/projects/tidy/.venv/bin/python /home/you/projects/tidy/organise.py >> /home/you/logs/tidy.log 2>&1
That runs at 07:00 every Monday. On Windows, Task Scheduler asks for the same three things through a wizard: create a basic task, set the trigger, then point the action at your interpreter, put the script path in the arguments box, and set "Start in" to the project folder.
Scheduled runs fail for boring reasons, and it is nearly always one of these four. The scheduler uses a minimal environment, so write absolute paths for everything, including the Python interpreter inside your virtual environment. The working directory is not your project folder, so a relative Path("orders.csv") resolves somewhere unexpected. Nobody sees your print output, so redirect it to a log file as shown above, or use the logging module. And the account running the task may not have permission to touch the folders you are targeting.
Test the exact command by pasting it into a terminal first. If it fails there, it will fail on schedule too, just silently.
Step 6: Turn the script into a reusable command
Hard-coded paths make a script single-purpose. argparse turns it into a tool you can point at anything, and it generates --help for free.
import argparse
from pathlib import Path
def main():
parser = argparse.ArgumentParser(description="Sort files into folders by type.")
parser.add_argument("source", type=Path, help="Folder to organise")
parser.add_argument("--apply", action="store_true",
help="Actually move files (default is a dry run)")
args = parser.parse_args()
if not args.source.is_dir():
parser.error(f"not a folder: {args.source}")
organise(args.source, dry_run=not args.apply)
return 0
if __name__ == "__main__":
raise SystemExit(main())
Making the dry run the default and requiring --apply for the destructive path is a deliberate choice. The safe thing should be the thing that happens when you type the command without thinking.
To call it as tidy from any directory, add a pyproject.toml next to your package and install it in editable mode.
[project]
name = "tidy"
version = "0.1.0"
requires-python = ">=3.11"
[project.scripts]
tidy = "tidy.cli:main"
[build-system]
requires = ["hatchling"]
build-backend = "hatchling.build"
Then run pip install -e . inside your virtual environment. The [project.scripts] entry means "create a command called tidy that calls the main function in tidy/cli.py". After that, tidy ~/Downloads --apply works from anywhere, and edits to the source take effect immediately because of the -e flag.
Mistakes that cost you an afternoon
No dry run. Every script that moves, renames or deletes should be able to describe what it would do without doing it. Build that in first, not after the accident.
Testing on real data. Copy a handful of files into a scratch folder and aim the script there. Rewriting the script is cheap. Recovering your invoices folder is not.
Swallowing errors. A bare except: around the whole loop hides the one file that broke and leaves you with a job that reports success while doing nothing. Catch the specific exception you expect, and let the rest crash loudly.
Growing forever. When a script passes a few hundred lines, split it into functions with clear inputs and outputs and add a few tests. Automation you cannot change is automation you will eventually abandon.
Where to go next
- Python Basics for the syntax and file-handling foundations this post builds on.
- Python tutorials for more scripting walkthroughs.
- Business Automation for deciding which processes are worth wiring together in the first place.
Stuck on a step, or does your script behave differently under the scheduler than in the terminal? Write to the desk and describe what you are seeing.
Comments
No comments yet. Be the first to share your thoughts.


