Advertisement
Open Source Projects by Phil Schwartz

Using DenyHosts Logs To Train A Simple Attack Detection Model

DenyHosts logs contain a useful record of hostile SSH activity: failed login attempts, repeated source addresses, blocked hosts, timestamps, and usernames targeted by attackers. With careful preparation, this operational data can become a small but practical machine learning dataset for detecting suspicious behavior before an automated block is triggered.

The goal is not to replace DenyHosts with a complicated artificial intelligence system. A lightweight classifier can support existing Linux security controls by assigning a risk score to new authentication events. This approach is especially useful for administrators who want transparent predictions, modest resource usage, and an audit trail they can inspect.

A successful project depends more on trustworthy labels and sensible features than on selecting an advanced algorithm. Logs must be normalized, duplicate events removed, and evaluation designed around time so the model is tested against behavior it has not already seen.

Why DenyHosts Logs Matter

SSH attacks often produce recognizable patterns. A single failed login may be harmless, while dozens of attempts against several usernames from one address within a short interval are considerably more suspicious. DenyHosts records can expose these patterns at a scale that is difficult to review manually.

The logs also reflect real operational decisions. An address added to a DenyHosts blacklist can serve as a useful positive label, while ordinary authentication failures may provide negative examples. That label is imperfect, since a host can be blocked for reasons unrelated to the exact event being classified, but it gives a practical starting point for supervised learning.

Log analysis should include surrounding server evidence when available. Apache activity, for example, can reveal whether an address is scanning other services or behaving like an automated crawler. Phil Schwartz’s guide to Scratchy log analysis offers a useful perspective on extracting behavioral signals from access logs.

Build A Reliable Training Dataset

Begin by collecting rotated DenyHosts files and related authentication logs from a controlled set of Linux systems. Preserve the original files, record their collection dates, and avoid mixing time zones. A parser should convert each line into structured fields such as event time, source IP, username, action, hostname, and message type.

The target variable must be defined explicitly. One simple design labels an event as malicious when its source IP is later added to the DenyHosts blocked list within a chosen period. Another labels aggregated IP windows rather than individual log lines, such as “this address generated suspicious activity during a 15-minute interval.” The second option often matches the real security decision more closely.

Be careful with class imbalance. Most authentication events may be routine, while only a small fraction lead to a block. Duplicating rare rows can distort results, so begin with class weights or carefully controlled undersampling. Split the data chronologically rather than randomly; otherwise, nearly identical attack bursts may appear in both training and test sets.

Turn Events Into Useful Features

Raw IP addresses are usually poor features. A model may memorize that a particular address was previously blocked, but that knowledge will not transfer to a new host in a changing attack campaign. Prefer behavioral and temporal attributes that describe what the source is doing.

Useful numeric features include failed attempts during the previous minute, five minutes, and hour; the number of distinct usernames targeted; the interval between attempts; and whether activity occurs across multiple services. You can also encode the hour of day, day of week, protocol, authentication method, and whether the address belongs to a known local network.

Categorical features need careful treatment. One-hot encoding works for a small set of message types or authentication methods, while IP reputation categories can be represented as external indicators. Avoid using future information, such as the final block status, in a feature calculated before prediction. That mistake creates data leakage and produces impressive but unusable accuracy.

Compare Simple Models

A transparent baseline is valuable. Start with a rule system, such as flagging an address after a threshold number of failures within a time window. Then compare it with logistic regression, a decision tree, and a random forest. These models are available through common Python libraries and can run comfortably on a small server.

Logistic regression provides readable coefficients: a high recent failure count may increase risk, while a long interval between attempts may reduce it. A shallow decision tree exposes threshold-like behavior that resembles administrator rules. A random forest can capture interactions, although its predictions are less straightforward to explain.

Approach Strength Limitation Suitable Use
Threshold rules Easy to audit and deploy Misses complex patterns Baseline protection
Logistic regression Fast and interpretable Assumes mostly linear effects Risk scoring
Decision tree Clear decision paths Can overfit noisy logs Explaining rules
Random forest Handles feature interactions Larger and less transparent Stronger offline model

The comparison should focus on operational value rather than a single headline score. A model that catches 95 percent of attacks but blocks many legitimate administrators may be worse than a conservative model with lower recall and far fewer false positives.

Evaluate Detection Without Fooling Yourself

Accuracy is rarely the right primary metric for intrusion detection. Precision shows how many alerts are genuinely useful, while recall measures how many labeled attacks are found. The F1 score balances these measures, but it may hide the cost difference between a missed attack and an incorrect block.

Use a time-based validation scheme. Train on earlier weeks, validate on a later period, and reserve the newest data for a final test. If several servers contribute logs, consider holding out one server entirely to discover whether the model generalizes beyond a single machine’s configuration.

Review false positives manually. Scheduled vulnerability scans, monitoring systems, backup jobs, and administrators with forgotten credentials can all resemble hostile activity. A confusion matrix, precision-recall curve, and list of the highest-risk events will reveal more than accuracy alone.

Practical Safeguards For Deployment

A prediction should initially create an alert rather than trigger an irreversible firewall action. Store the model score, contributing features, source address, and eventual administrator decision. This feedback can improve labels and provide evidence when a legitimate user is blocked.

Keep the existing DenyHosts threshold as a safety boundary while the classifier operates in shadow mode. Retrain periodically only after reviewing recent errors, and version both the model and feature-generation code. Security data changes as attackers change tactics, so a model that performs well this month may need recalibration later.

Move From Experiment To Protection

A small Python pipeline can handle the complete workflow: parse DenyHosts records, maintain rolling counters, calculate features, load a serialized classifier, and write a risk event for review. Keeping each stage separate makes it easier to replace the model without rewriting the log collector or response system.

The most useful result may be a ranked stream of suspicious sources rather than a binary verdict. Administrators can combine that score with existing DenyHosts decisions, SSH configuration, firewall rules, and evidence from other services. This layered design preserves the reliability of established defenses while adding statistical context.

Begin with a week of historical logs, document the labeling rule, and measure false positives before allowing automated responses. Once the model proves stable on future data, integrate its scores into monitoring and incident workflows, giving the system a practical path from offline experiment to dependable Linux security tool.