Advertisement
Open Source Projects by Phil Schwartz

File locking with Python fcntl in multi-process systems

When several Python processes need to touch the same file at the same time, the outcome is rarely pretty. A worker in Brisbane might be appending a JSON record while a scheduler in Sydney reads it for indexing, and without coordination both end up with corrupted bytes. File locking solves that, and on Linux and macOS the most direct path is the fcntl module, which wraps the POSIX advisory locking primitives.

The fcntl module is thin, fast, and built into CPython. It exposes flock, a function that hands out read locks, write locks, and unlock operations directly on a file handle. Because the locks are advisory, every cooperating process must opt in, but in a closed ecosystem of Python daemons this is rarely a problem. This piece walks through how to use it, where it shines, and where you need to be careful, with a few practical patterns drawn from Australian production environments.

Why multi-process coordination matters

Multi-processing remains the default path for CPU-bound Python workloads. The GIL limits threading, so when a team in Melbourne needs to crunch log analysis or a Sydney fintech runs parallel risk calculations, they reach for multiprocessing. Those workers usually need to read or update a shared file at least once: a configuration cache, a progress marker, a rotating log, or a lockfile acting as a sentinel.

Australian privacy expectations make coordination especially important. Under the Notifiable Data Breaches scheme, organisations that mishandle personal information can face reporting obligations and penalties, so a corrupted audit log caused by racing writers is more than a nuisance. Even outside regulated industries, software shops in Adelaide and Perth that build tools for the resources sector often run daemons across many machines, and a single shared state file can become a coordination choke point.

How fcntl.flock works

At its core, flock operates on an open file descriptor. You call it with two arguments: the file descriptor and an operation. The operation is a bitwise combination of a lock type and optional flags. LOCK_SH grants a shared lock suitable for readers, LOCK_EX reserves exclusive access for writers, and LOCK_UN releases the lock. Adding LOCK_NB makes the call non-blocking, so a process that cannot acquire the lock raises BlockingIOError immediately instead of stalling.

There are two flavours of POSIX locking: flock, which works on file descriptors and treats each open file description as its own lock target, and fcntl-style locks based on the F_SETLK command, which are associated with the underlying file and the process. The fcntl module exposes both, but for most Python work flock is the friendlier choice because the semantics are simpler and the file handle can be a context manager that automatically closes the descriptor. The downside is portability: flock behaves differently across NFS mounts and is unavailable on Windows, which matters when a developer in Hobart tries to run the same script on a Windows build agent.

Building a reusable lock wrapper

The simplest pattern wraps flock in a context manager. Open the target file, call flock with the chosen mode, and let exit handle unlocking and closing. This makes locking declarative at the call site, so a worker function reads almost like synchronous code.

import fcntl
import contextlib

@contextlib.contextmanager
def file_lock(path, mode=fcntl.LOCK_EX):
    f = open(path, 'w')
    try:
        fcntl.flock(f, mode)
        yield f
    finally:
        fcntl.flock(f, fcntl.LOCK_UN)
        f.close()

Points to keep in mind before adopting this pattern:

Pitfalls and real-world edge cases

Cross-filesystem behaviour trips up many teams. Locking on an NFS share, common in older data centres in regional Victoria, can return success even though the remote server does not honour the request. The result is silent corruption that only shows up under contention. Local filesystems on ext4, XFS, and APFS handle flock reliably, so the rule is to keep shared state on local disk or on a cluster-aware filesystem like OCFS2.

Stale locks after a crash deserve special attention. Because locks are tied to file descriptions and processes, when a Python process is killed with SIGKILL its locks disappear with the kernel structures. There is no leftover lockfile to scrub, which is one of flock's quiet strengths over traditional /var/lock/subsystem/ lockfiles. The trade-off is that there is also no visible state to inspect, so adding structured logging around lock acquisition makes incidents in production easier to diagnose.

Patterns that hold up in production

Once the basics are in place, a few patterns keep showing up in code reviews around Sydney and Perth engineering teams. The first is using a dedicated lockfile rather than locking the data file itself. A tiny sentinel file, often a hidden dotfile alongside the data, lets readers and writers coordinate without having to open the data file in a conflicting mode.

The second is layering fcntl with multiprocessing.Event for intra-process signalling. The OS lock keeps processes apart, while the Event speeds up workers on the same node that would otherwise retry blindly. The third is wrapping the lock in a tiny class so logging, metrics, and timeouts become uniform across the codebase. Notes on that class, along with a couple of related utilities, appear in the ricoblog.net archive.

Practical habits that travel well between teams:

Teams shipping software that runs across Linux distributions used by universities, government departments, and small SaaS startups across Australia tend to keep their locking layer boring and small. Anything fancier, such as Redis-backed distributed locks, belongs only when the workload actually spans machines, which is rarely the case for a single-host daemon.