Advertisement
Open Source Projects by Phil Schwartz

Designing DenyHosts’ data model for attack sources

When I designed DenyHosts, I wanted it to solve a focused problem: identify repeated SSH attacks and prevent the same sources from continuing to hammer a Linux machine. The program needed to run quietly in the background, use very few resources, and remain understandable to administrators who might need to inspect or repair its files by hand.

That requirement shaped the data model from the beginning. Instead of introducing a database server, I used simple text files to record source addresses, suspicious activity, successful logins, and hosts that should be blocked. The model was small, portable, and appropriate for machines that might have limited memory, slow disks, or no additional services installed.

An attack source is more useful than an isolated failed login. A single authentication error can result from a mistyped password, a misconfigured script, or a legitimate user connecting from the wrong account. Repeated failures from the same IP address provide stronger evidence, so DenyHosts had to aggregate events rather than treat every log line as an independent decision.

That approach remains relevant for Australian operators. A small business in Perth, a developer running a server in Melbourne, or a hosting customer in Sydney may receive constant automated scans even when the machine is not publicly advertised. The data model had to make a cautious decision quickly while avoiding unnecessary disruption to legitimate users on dynamic NBN connections.

Starting with the source address

The central record in DenyHosts is the host address associated with suspicious SSH activity. In the original design, this generally meant an IPv4 address extracted from authentication logs. The address became the key used to combine multiple observations from different failed-login events.

Using the source address as the primary identity made the implementation practical. A log parser could find an address, update the relevant counters, and compare the accumulated evidence with configured thresholds. There was no need to maintain a complex relationship between users, sessions, processes, and individual log entries.

The source itself was not treated as proof of malicious intent. It was an identifier around which evidence could be grouped. That distinction matters because addresses can represent a shared office network, a university connection, a cloud provider, or a carrier-grade NAT pool. Blocking an address can affect several people, so the model had to support thresholds and exceptions rather than an immediate reaction to one event.

Separating evidence from enforcement

I separated the information that described an attack from the information that controlled access. Files such as the suspicious-login records captured failed attempts by source and username, while denied-host records represented the stronger decision to block an address.

This separation gave DenyHosts a useful progression. A host could first appear as an observed source, then accumulate enough failures to become suspicious, and finally enter the denied list. A successful login could be recorded separately, allowing the administrator to see that an address had authenticated correctly even if it had also generated failed attempts.

The model also supported distinct categories for attacks against root, invalid users, and ordinary accounts. That categorisation helped preserve context without requiring a large event database. It let configuration decide whether, for example, repeated attempts against nonexistent accounts should count differently from failures involving a real user.

Keeping storage plain and inspectable

Flat files were a deliberate design choice rather than an accidental limitation. A system administrator could open them with standard Unix tools, copy them during an incident, back them up, or remove a corrupt entry without learning a database query language. This was especially valuable for small servers managed remotely.

The files acted like compact tables with stable, recognisable purposes. One stored addresses that had been denied, another tracked suspicious activity, and others recorded special classes of login failures or trusted hosts. The design avoided duplicating every raw log line, which kept storage requirements modest while preserving the counts needed for decisions.

This style also suited the open-source environment around DenyHosts. Python provided straightforward file handling, and Linux distributions could package the program without adding a database dependency. A volunteer administrator in Adelaide or a small web shop in Brisbane could install it alongside OpenSSH and inspect its state using ordinary shell commands.

Handling time and changing behaviour

Attack evidence has a lifespan. A source that generated thousands of failures last year may no longer be relevant, while an address that has just started scanning the server deserves attention now. DenyHosts therefore needed time-aware records and a way to expire old information.

The data model supported purge operations so that historical entries did not grow without limit. It also tracked processing progress through an offset, allowing the daemon to resume reading a log from the correct position instead of scanning the entire file after every restart. That reduced repeated work and made recovery more predictable.

Time also affects Australian deployments. A server in Darwin can operate on a different local clock from an administrator in Hobart, while daylight-saving changes affect parts of the country but not others. For security records, consistent timestamps and careful log parsing are more important than displaying dates in the operator’s local format.

Protecting trusted and legitimate hosts

A blocking system needs an explicit way to express trust. DenyHosts included allowed-host data so that an administrator could protect a known address or network from automatic denial. This was essential for avoiding a lockout when a user made repeated mistakes from an important management connection.

The exception model was especially useful where addresses changed regularly. Australian households and small offices often receive dynamic service addresses, and mobile or remote workers may appear from different networks. The practical answer was to maintain a trusted range carefully, use secure administrative access, and avoid assuming that every repeated failure represented an attacker.

The same principle applies to commercial infrastructure. A managed service provider, monitoring platform, or backup system may legitimately connect many times from a shared address. Recording evidence separately from the deny decision meant those arrangements could be reviewed and exempted without erasing the underlying observations.

Making updates safe across processes

A daemon that reads logs while another process examines or purges its records can corrupt its own state unless file access is coordinated. DenyHosts used lock files and controlled update operations so that two processes would not write conflicting versions of the same data at once.

The update sequence was designed around small, recoverable changes. Read the current state, modify the relevant source record, write the new information, and release the lock. This approach reduced the chance that a restart, scheduled purge, or manual administrative action would leave a half-written deny list.

The final model was intentionally conservative. It stored enough structure to count attacks, distinguish categories, remember trusted and denied sources, and continue safely after interruption. It did not attempt to become a full security information and event management platform. For DenyHosts, a clear source-oriented model was the right balance between useful protection, low overhead, and practical administration.