Building a Python Filename Sanitizer for Windows, macOS and Linux
Cross-platform development is something Australian engineers grapple with constantly. A script written on a MacBook in Surry Hills might end up running on a Windows machine in a Brisbane government department, or on a Linux box hosted at a data centre in Sydney's Global Switch facility. When that script produces output files, the names it picks have to survive the trip, and that is where a robust filename sanitizer earns its keep.
File naming conventions are deceptively simple. Most of us treat filenames as plain strings until something breaks: a colon in a filename that Linux tolerates but Windows rejects, a trailing space that macOS cheerfully accepts but Windows refuses to copy, or a slash that one platform interprets as a path separator and another treats as a perfectly legal character. The Australian Cyber Security Centre publishes guidance on safe handling of files, and while the focus is security rather than portability, the principle of not letting ambiguous input reach storage applies equally well to filenames.
A reliable sanitizer does more than strip out a hard-coded list of banned characters. It needs to behave predictably on every supported platform, handle Unicode sensibly, and produce names that are still recognisable to the people who will eventually open them. Anyone who has watched a colleague in Adelaide hunt through a folder full of files called "doc__final__FINAL_v2.pdf" knows that the human side of the problem matters too.
In practice, the same function often sits alongside log parsers and other utilities in a shared internal toolkit, the kind of small library that quietly does the boring job of keeping everything else from breaking. The approach below leans on standard library modules so it can be dropped into any project without adding new dependencies.
Knowing What Each Filesystem Forbids
Every mainstream operating system draws its banned-character list from somewhere different. Windows reserves a set of reserved characters that dates back to MS-DOS, including angle brackets, colon, double quote, pipe, question mark and asterisk. It also blocks reserved device names such as CON, PRN and AUX, a quirk that still trips up developers who assume the rules will match POSIX. macOS, by contrast, treats the colon as a legacy path separator inherited from Classic Mac OS, and modern versions also flag the slash. Linux distributions generally allow almost anything except the slash and the null byte, which makes them permissive rather than chaotic.
Understanding these rules matters because a sanitizer that targets only Windows often lets through filenames that confuse shell scripts, while one that targets only POSIX systems risks creating files that fail to copy onto an attached NTFS drive. A practical implementation usually starts with a broad set of forbidden characters and then narrows it for the target filesystem. Unicode normalisation also deserves attention: a name that looks identical to a human eye may compare as different bytes, which is exactly the kind of surprise you do not want when comparing files across machines.
The Australian government's Digital Transformation Agency has pushed federal departments towards interoperable document standards, which indirectly encourages developers working on government contracts to think about how their files will move between Mac workstations in Canberra and Windows servers managed elsewhere. A sanitizer that is explicit about its policy is far easier to audit than one that relies on trial and error.
Handling Collisions Without Losing the Original
Stripping dangerous characters is only the first half of the job. Once the name is sanitised, the function still has to deal with the fact that two different inputs might produce the same safe output. A folder full of reports called "report.pdf", "report (1).pdf", "report (2).pdf" is functional, but it is also annoying for anyone trying to find a specific file, and it is the sort of friction that prompts users in Perth or Hobart to email support rather than dig through the mess.
A common technique is to add a short hash suffix when a collision occurs, derived from the original input. This preserves uniqueness while keeping the readable prefix intact, so "Quarterly Update: Sales.pdf" might become "Quarterly Update Sales_a3f9.pdf" rather than colliding with another file that stripped down to the same string. The hash should be deterministic, since two runs of the script on different days will both need to produce the same name for the same input.
It is worth thinking about the order of operations here. Sanitise first, then check for collisions, otherwise the comparison logic ends up working on unsafe strings and you have to repeat the work. In a codebase that already deals with logging, an approach similar to the one described in parsing Apache logs keeps the helper functions consistent and easier to maintain across the project.
Keeping Names Within Filesystem Limits
Maximum path length is the third pitfall, and it varies dramatically between platforms. Older Windows releases capped paths at 260 characters through the MAX_PATH limit, although modern versions can opt into longer values with the right manifest. macOS's HFS+ had a 255 UTF-16 code unit limit, and APFS keeps that limit but with cleaner semantics. Most Linux filesystems top out at 255 bytes for the filename component, which interacts awkwardly with multibyte characters.
A good sanitizer truncates the name before the extension, leaving the suffix intact. Trimming in the middle of a word feels awkward, so a softer fallback is to chop at the nearest word boundary. The character count should also account for the eventual collision suffix and any path prefix the caller might add. For Australian teams writing files that may end up on SharePoint Online, which adds its own URL-encoding quirks, it pays to be conservative. Truncating to 120 characters for the filename component leaves comfortable headroom for both a hash suffix and any folder hierarchy.
Preserving Readability for the End User
Stripping punctuation aggressively turns "Brisbane Meeting - Friday (notes).md" into "Brisbane Meeting Friday notes.md", which is fine. Stripping it too aggressively turns it into "BrisbaneMeetingFridaynotesmd", which is awful. The art is to replace banned characters with sensible substitutes rather than deleting them, so a colon becomes a hyphen, a slash becomes a dash, and curly quotes become straight quotes.
Diacritics deserve a decision. For a script running in a multilingual environment, keeping them in place makes sense. For a workflow that feeds into systems with poor Unicode support, ASCII folding may be the safer choice. The Privacy Act 1988 and the Notifiable Data Breaches scheme shape how Australian organisations handle personal information in files, which is a separate but related concern: a sanitizer that quietly mangles a person's name in a filename can make compliance reviews harder than they need to be.
Putting the Function to Work
Once the sanitizer is written, it slots naturally into any pipeline that writes files based on user input or external data. Web scrapers, log rotators, report generators and file upload handlers are all candidates, and the same helper covers each case. Wrapping the function in a small module with a clear docstring and a handful of unit tests is usually enough to keep it dependable.
For teams that already maintain an internal toolkit, the new function joins the existing utilities without ceremony. Adding it to the test suite alongside log parsers and other shared helpers means regressions get caught early, and the project stays predictable for the next developer who picks it up.
