Advertisement
Open Source Projects by Phil Schwartz

Cleaning Up DenyHosts Without Breaking Its Purpose

DenyHosts began with a focused objective: watch SSH authentication failures, identify abusive addresses, and prevent repeated attacks from consuming a server’s attention. The first versions achieved that goal with relatively little code, which made the project easy to install and practical for administrators running modest Linux systems.

That simplicity also hid a growing maintenance cost. As features accumulated, early shortcuts became technical debt: state was stored in formats that were difficult to evolve, security decisions were mixed with file handling, and tests covered successful execution more thoroughly than failure recovery. The program worked, but its internal design made every change riskier.

Cleaning up that debt required more than rewriting functions. I had to preserve existing behavior for long-time users, understand how installations differed in the field, and improve the architecture without turning a small defensive utility into an elaborate framework.

Where The Debt Began

The earliest implementation reflected the environment in which it was written. A few flat files, regular expressions, and shell-oriented assumptions were enough to record failed logins and update a deny list. That approach was efficient for a prototype, but it created hidden coupling between log parsing, host tracking, configuration, and command-line behavior.

One problem was that data formats became accidental interfaces. Administrators inspected files manually, scripts relied on their layout, and later code assumed that every line was complete and correctly ordered. A malformed entry or interrupted write could therefore affect more than one part of the application.

The code also favored immediate results over explicit boundaries. A parser could update state while it was still interpreting a log line, and a storage routine could make policy decisions because the two responsibilities lived close together. Those shortcuts reduced the number of files initially, yet increased the number of assumptions I needed to preserve.

Separating Security Logic From State

The first major repair was to divide the system into clearer layers. Log readers became responsible for finding authentication events. A decision component determined whether an address had crossed a threshold. Persistence code stored observations and blocked hosts, while the command-line layer handled options and reporting.

This separation made security policy easier to audit. A change to the number of allowed failures no longer required understanding how a particular log format was written to disk. It also made alternate log sources less dangerous to add, since each parser could produce the same internal event representation.

I treated compatibility as a design constraint rather than an inconvenience. Existing configuration names and common file locations had to continue working, but new code could normalize values at the boundary. Internally, the program could then operate on consistent types instead of repeatedly interpreting strings and legacy defaults.

Making Persistence Predictable

File handling became the next focus. DenyHosts runs in environments where processes can be interrupted, permissions can be unusual, and several service instances may compete for the same state. A direct write to a shared file was too fragile for a security tool whose decisions depend on reliable history.

I introduced more disciplined update behavior: validate data before writing it, use temporary files where appropriate, and replace completed files atomically. Locking and clearer error reporting reduced the chance that a partially written deny list would silently become the new source of truth. The goal was not theoretical perfection; it was graceful recovery from ordinary operational failures.

I also stopped treating every missing file as the same event. A first run, an unreadable file, an empty file, and a corrupted file require different responses. Distinguishing those cases helped administrators diagnose installations and prevented the software from quietly discarding useful evidence.

Area Early implementation Repaired approach
Log processing Parsing and state updates were intertwined Parsers emit normalized events
Configuration Values were interpreted throughout the code Inputs are validated at the boundary
State files Direct writes and implicit formats Validated, safer replacement and recovery
Error handling Failures could look like empty data Errors are classified and reported
Testing Success paths received most attention Recovery, malformed input, and compatibility are tested

Testing The Failure Paths

The most valuable tests were not the ones proving that a normal SSH failure could be detected. They covered truncated lines, changing timestamps, unexpected whitespace, duplicate addresses, rotated logs, inaccessible directories, and state files containing old or invalid records.

I built fixtures around real operational scenarios rather than idealized function calls. A test could create a small log, run the parser, interrupt a write, and verify that the next execution recovered without losing the existing deny list. This made regressions visible at the level users actually experienced them.

Small developer tools also influenced how I approached parsing and debugging. When regular expressions became difficult to reason about, I used the Kodos regex debugger to inspect patterns and edge cases before embedding them in production logic. Better tooling did not replace tests, but it shortened the path from a confusing input to an explainable fix.

Reducing Operational Friction

Technical debt is often measured in code complexity, but administrators experience it as friction. Ambiguous messages, undocumented defaults, inconsistent exit statuses, and unclear permissions can make a sound security policy feel unreliable. I treated those details as part of the software’s interface.

Documentation became more explicit about installation paths, log formats, service behavior, and the meaning of each persistent file. Command output was revised to distinguish an address that was newly blocked from one that was already known. Those changes reduced support burden and made automated administration safer.

I also reviewed licensing and distribution assumptions as the project matured. Open-source users need to know what they can modify, redistribute, and integrate into their own environments. A maintainable project includes that information alongside its code, downloads, and release history rather than leaving it scattered across old notes.

Practices That Keep Debt Small

The cleanup changed how I evaluate new features. A request may sound small, but it can introduce another implicit file format, another global setting, or another path that bypasses validation. Before accepting a change, I look for the boundary it belongs behind and the failure mode it creates.

The following practices have been especially useful for keeping a compact security utility understandable:

The result is not a perfect or frozen codebase. DenyHosts still has to operate across different Linux distributions, Python environments, log conventions, and administrative habits. The difference is that new changes now have clearer places to live, and failures are more likely to be visible, recoverable, and testable.

Early shortcuts were useful because they helped the project exist. The important lesson was not to regret them, but to recognize when they had stopped serving the project. By replacing hidden assumptions with explicit interfaces and dependable state handling, I could preserve DenyHosts’ practical character while giving future maintainers a safer foundation.

Explore the DenyHosts project, its open-source history, and the related development utilities to see how small Linux tools can evolve without losing their usefulness.