Making DenyHosts Blocklist Sync Race-Free
DenyHosts was designed to protect SSH services by detecting repeated login failures and adding hostile addresses to a local deny list. Its synchronization feature extended that protection across multiple machines, allowing one host to share attack data with others. That shared state introduced a subtle concurrency problem: two processes could modify the same blocklist at nearly the same time.
The failure was intermittent, which made it especially difficult to reproduce. Most sync runs completed normally, but under the right timing, a newly discovered address disappeared, a downloaded update overwrote a local change, or a partially written file was read by another process. The system looked reliable until several workers, timers, or administrator commands overlapped.
The fix required more than adding a sleep or retry loop. I had to define which operations were allowed to run concurrently, protect the critical sections, and make file updates atomic. The experience also reinforced a useful lesson from maintaining small Linux utilities: simple text files still need transaction-like handling when multiple actors can write them.
Finding The Lost Update
The first clue came from comparing event logs with the final contents of the synchronized blocklist. A host recorded a failed SSH attempt and reported the address, but another machine never retained it after a synchronization cycle. In a separate case, the receiving host contained an older list even though the transfer log showed that newer data had arrived.
The common sequence was a read-modify-write operation. A local process read the current deny list into memory, while the sync worker read a remote copy. Each process then produced its own merged result and wrote it back. Whichever write happened last replaced the other result, even if that process had started with stale data.
This was a classic lost-update race rather than a networking failure. The synchronization protocol could transfer valid data, but the local persistence layer had no protection against concurrent writers.
Reproducing The Timing Window
Intermittent bugs need deliberate pressure. I created a test harness that launched local updates and synchronization jobs together, then inserted short delays between reading the file, merging entries, and writing the result. Repeating that cycle exposed missing addresses within a few runs.
I also tested unusual but realistic conditions: an empty file, duplicate addresses, a process interrupted during writing, and a remote list that arrived while a local scan was recording new failures. Logging process identifiers and timestamps made the ordering visible. Without that detail, the output only showed that the final list was wrong, not why it became wrong.
The investigation benefited from the same evidence-first habit I use when examining web server activity. My work on Scratchy’s design story had already taught me how much useful information can be hidden in carefully structured logs.
Separating Coordination From Data Merging
The first design decision was to make synchronization a serialized operation. A process-level lock prevented two DenyHosts workers on the same machine from entering the blocklist update routine simultaneously. The lock covered the complete read, merge, and commit sequence rather than only the final write.
That boundary mattered. Locking just the write call would still allow two workers to read the same old state and calculate conflicting results. By holding the lock from the initial read through the final replacement, every worker merged against the latest committed version.
The merge itself remained idempotent. Addresses were treated as a set, so receiving the same report twice did not create duplicates. Local discoveries and remote entries were combined before the result was saved, preserving both sources of information.
| Operation | Risk Before The Fix | Protection After The Fix |
|---|---|---|
| Read local blocklist | Read while another process was writing | Performed under the update lock |
| Receive remote entries | Remote data could replace local discoveries | Merged with the current local set |
| Write updated list | Partial or stale output was possible | Written to a temporary file first |
| Replace active file | Readers could see incomplete content | Atomic rename used for commit |
| Interrupted update | Corrupt state could remain | Original file stayed intact until commit |
Making File Writes Atomic
A lock protects cooperating processes, but it cannot guarantee that a reader will never encounter an incomplete file. An administrator, monitoring script, or older process might open the blocklist without taking the same lock. I therefore changed the write path to use a temporary file in the same directory.
The complete merged content was written and flushed to that temporary path. Once the write succeeded, the temporary file was renamed over the destination. On the target Linux filesystems, rename provided the atomic transition needed here: readers saw either the old complete file or the new complete file, rather than a half-written stream.
Using the same directory was important because atomic replacement depends on filesystem semantics. A temporary file in another location could require a copy operation, which would reintroduce the partial-write problem. Cleanup logic also removed abandoned temporary files after failures.
Handling Remote State Safely
The synchronization protocol needed a clear rule for conflicts. A remote list was treated as additional evidence, not as an authoritative replacement for local state. That choice prevented a stale peer from erasing a recent local denial.
After downloading remote entries, the worker acquired the local lock, reread the active blocklist, and merged the remote data with that freshest local snapshot. This second read was intentional. Reading before network activity would leave a long window in which another process could update the file.
The worker then committed the merged set and updated its synchronization metadata as part of the same protected workflow. If the download or merge failed, the existing blocklist remained available and the next run could retry without silently advancing its state.
Testing Failure And Recovery Paths
Normal success cases were not enough. I added tests for simultaneous writers, duplicate reports, interrupted temporary-file writes, malformed remote records, and a worker that lost its lock before committing. The important assertion was that a failed operation could not remove an address already present in the active list.
I also tested repeated synchronization because a safe operation should converge. Running the same remote update multiple times had to produce the same content, with stable ordering where practical. Deterministic output made both regression tests and administrator reviews easier.
The final tests exercised real filesystem behavior rather than relying entirely on mocks. Concurrency bugs often hide at the boundary between application logic and operating-system primitives, so lock acquisition, rename behavior, permissions, and cleanup all deserved coverage.
Practices That Prevent A Repeat
The repair was small in code but broad in its implications. A blocklist is security-sensitive state, even when it is stored as plain text. Every path that can modify it must follow the same coordination and commit rules.
The most durable safeguards were these:
- Use one shared lock for every local blocklist mutation.
- Re-read the active file immediately before merging remote entries.
- Write complete content to a same-directory temporary file.
- Replace the destination with an atomic rename after a successful write.
- Keep synchronization metadata consistent with the blocklist commit.
- Test overlapping workers and interrupted operations as first-class cases.
This approach made DenyHosts synchronization predictable under load and easier to reason about during maintenance. It also turned an elusive production symptom into a set of explicit guarantees: no lost local discoveries, no partial blocklists, and no silent replacement by stale remote data.
Reliable open-source tools are built through this kind of careful refinement. Explore the DenyHosts project, review its implementation, and use these concurrency patterns when a security list or configuration file must be shared safely across processes.
