Using Scratchy to Detect Distributed Scanning Attacks Across Multiple Subnets
A single hostile IP address is easy to notice in an Apache access log. A distributed scan is harder: dozens or hundreds of systems may probe the same application, exploit path, or administrative endpoint while staying below the alert threshold for any individual address. When those systems are spread across multiple subnets, ordinary per-host reporting can hide the wider pattern.
Scratchy, the Apache log analyzer in Phil Schwartz’s collection of open-source software, provides a practical way to turn raw web requests into searchable evidence. Its reports can help administrators identify repeated probes, unusual status-code patterns, suspicious user agents, and bursts of requests that become significant only after several network ranges are considered together.
The tool is most effective when treated as an investigation aid rather than a complete intrusion detection system. It can reveal relationships in historical access data, helping security teams decide which addresses, paths, and time windows deserve deeper analysis in firewall logs, reverse-proxy records, or host monitoring systems.
What Distributed Scanning Looks Like In Web Logs
Distributed scanning is an activity pattern in which multiple source addresses perform related reconnaissance. The requests may target common files such as login pages, environment files, version-control directories, backup archives, or known framework endpoints. Each address might generate only a few requests, but the combined activity forms a recognizable campaign.
Across several subnets, the source addresses may belong to different cloud providers, compromised hosts, residential networks, or rotating proxy services. Their geographic and network diversity can make a simple “top clients” report less useful. The important signal is often the shared behavior: identical URI probes, matching user-agent strings, similar request intervals, or repeated responses such as HTTP 404, 403, and 401.
Scratchy can help expose these relationships by summarizing Apache access logs by request, response, client, and time period. Exported or filtered results can then be grouped by subnet so that a low-volume source is evaluated in the context of all similar sources.
Preparing Apache Logs For Reliable Analysis
Before analyzing a suspected scan, verify that the access logs contain the fields needed for correlation. The client address, timestamp, request method, URI, protocol, status code, response size, referrer, and user-agent string are especially useful. Consistent timestamps matter when requests are distributed across servers or subnets.
Centralize copies of logs from the relevant web servers when possible. If each server is analyzed independently, a scan that moves between applications may appear fragmented. Preserving the original files and documenting the timezone also prevents investigators from mistaking normal traffic bursts for coordinated activity.
Normalize network information outside the raw log when necessary. IPv4 and IPv6 addresses should be handled consistently, and internal reverse proxies must be configured so that the recorded client address represents the actual source rather than the proxy itself. Scratchy’s output becomes much more valuable when the input accurately reflects the request path.
Turning Scratchy Reports Into Correlation Evidence
Begin with a broad report covering the suspected time range. Look for high-frequency paths, unusual status codes, repeated methods, and clients that request many nonexistent resources. A scan may be visible through a concentration of failed requests rather than a high total request count.
Next, compare the suspicious URIs against the source addresses. If many unrelated clients request the same unusual path within a short interval, that shared target is stronger evidence than any single address ranking. The same method works for user agents, HTTP versions, referrers, and response patterns.
The investigation should then move from individual IPs to network groups. Map each source to its subnet or autonomous system, and count distinct sources, requested paths, and time intervals within each group. A campaign may involve a small number of requests from each subnet while producing a large combined footprint.
| Signal | What Scratchy Can Reveal | Why It Matters |
|---|---|---|
| Repeated unusual URIs | Probes for administration, backups, or known vulnerabilities | Shows shared reconnaissance objectives |
| HTTP 404 and 403 bursts | Requests for files or directories absent from the application | Helps distinguish scanning from normal browsing |
| Common user agents | Repeated automated client identification | Links otherwise unrelated source addresses |
| Short synchronized time windows | Similar activity across several hosts | Supports a coordinated-campaign hypothesis |
| Subnet concentration | Multiple clients within related network ranges | Helps prioritize blocking and investigation |
| Low-volume distributed requests | Small contributions from many addresses | Prevents per-IP thresholds from hiding the event |
Comparing Sources Across Multiple Subnets
Subnet-level aggregation should be careful and specific. Grouping by a broad prefix can combine unrelated customers, while grouping too narrowly can obscure a coordinated source pool. Begin with the organization’s own network boundaries, then compare the results with public registration data or trusted threat-intelligence sources.
A useful finding might be that six addresses from three cloud subnets requested the same set of sensitive paths within fifteen minutes. Another pattern could involve separate residential ranges using the same malformed query string. These observations do not prove that every source belongs to one operator, but they provide a defensible basis for escalation.
Compare suspicious traffic with a baseline from normal business periods. Search engines, monitoring services, vulnerability scanners, and mobile clients can generate unusual requests without being malicious. Frequency, timing, requested resources, and response outcomes should be evaluated together before access controls are changed.
Reducing False Positives During Investigation
Automated log analysis can overstate the importance of a pattern when an application or deployment process creates unusual traffic. Health checks may request a fixed endpoint repeatedly, while a broken frontend can cause legitimate users to generate many 404 responses. Known scanners used by the organization should be labeled before review begins.
Prioritize combinations of indicators instead of single fields. A request for a nonexistent file is weak evidence by itself. The same request becomes more significant when it is repeated by many addresses, paired with several exploit-related paths, and concentrated in a narrow time window.
Keep a record of the evidence that led to each decision. Save the relevant Scratchy output, source log lines, subnet mappings, and timestamps. This makes it easier to compare later activity and prevents a temporary block from becoming a permanent rule based on incomplete information.
Acting On Findings Without Losing Visibility
Once a distributed scan is confirmed as suspicious, use the analysis to guide layered controls. Rate limiting, web application firewall rules, provider-level abuse reports, and targeted firewall blocks can reduce exposure. Avoid relying solely on a large static denylist, since distributed activity often changes addresses quickly.
DenyHosts and similar host-protection tools are more closely associated with services such as SSH, while Scratchy focuses on web access records. Used together, they can provide broader visibility: Scratchy identifies application-layer reconnaissance, and host-based controls respond to repeated authentication abuse. The tools should remain part of a wider monitoring and incident-response process.
Useful operational recommendations include:
- Retain rotated Apache logs long enough to compare recurring scan waves.
- Aggregate findings by URI, user agent, timestamp, subnet, and autonomous system.
- Establish normal traffic baselines for each public application.
- Alert on coordinated low-volume probes, not just the busiest individual IPs.
- Preserve raw evidence before applying blocks or rewriting log filters.
Scratchy is particularly valuable after an event, when investigators need to understand how widely a scan spread and whether several applications were targeted. Its reports can also support preventive tuning by showing which paths attract repeated automated attention.
Use Scratchy to examine a representative set of Apache logs, correlate suspicious requests across your network ranges, and preserve the resulting evidence alongside your security records. By shifting attention from isolated IP addresses to shared behavior, you can make distributed reconnaissance easier to detect and respond to before it develops into an application compromise.
