Tracing Coordinated Intrusions with Scratchy Forensic Analysis
When a web server starts returning unusual traffic patterns, the first challenge is separating a genuine attack from the constant noise of the public internet. A single failed request can be harmless scanning, a crawler misconfiguration, or the first visible trace of a larger campaign. The useful evidence often appears only when thousands of log lines are examined together.
I used Scratchy, my Apache log analyzer, to investigate this kind of activity. Rather than treating each source IP as an isolated problem, I analyzed request timing, URLs, response codes, user-agent strings, and access frequency as connected signals. That shift made it possible to identify coordinated behavior that was distributed across many addresses.
The investigation also reinforced a principle that applies to most Linux security work: logs are more valuable when they are preserved, normalized, and compared systematically. Scratchy provided the practical starting point, while manual review and lightweight scripting helped test the conclusions.
Starting With Raw Apache Logs
The first step was to preserve the original access logs before filtering or rotating them. I copied the relevant files into a restricted analysis directory, recorded their timestamps, and worked from those copies. Keeping an untouched source is important because an early assumption can otherwise remove evidence needed later.
Apache logs contain several useful fields: client address, timestamp, request method, path, protocol, status code, response size, referrer, and user-agent. Scratchy made it easier to summarize these fields without losing the relationship between individual requests. I could quickly see which paths were being targeted and whether the server was responding with successful pages, redirects, missing-file errors, or access denials.
The initial report showed what looked like ordinary internet scanning. There were many IP addresses, a mixture of probes for administrative paths, and repeated requests for files that did not exist. The distribution across addresses initially made the activity appear random, but the request sequence contained a more consistent pattern.
Finding the Shared Behavior
I grouped requests by URL and then examined their timestamps. The same small set of paths appeared across unrelated addresses within narrow time windows. The clients did not always use identical headers, but their behavior was similar enough to suggest a common script or coordinated instruction set.
The strongest signal was the order of operations. Each source tended to request a landing page, follow it with probes for known administrative locations, and then attempt several filenames associated with vulnerable applications. That sequence repeated even when the source addresses belonged to different networks.
User-agent strings provided supporting evidence, although they were not treated as proof. Some requests used common browser identifiers, while others advertised command-line tools or generic crawlers. Attackers can change these values easily, so I used them as a correlation field rather than as an identity marker.
Comparing Signals Across Requests
A forensic review becomes clearer when each indicator is given appropriate weight. Source IP addresses are useful for immediate blocking, but they are weak evidence of campaign identity because attackers can use proxies, cloud hosts, compromised systems, or rotating infrastructure.
| Signal | What It Revealed | Reliability |
|---|---|---|
| Request timing | Synchronized bursts from multiple clients | High |
| Requested paths | Repeated probe sequence | High |
| Status codes | Whether targets existed or were rejected | Medium |
| User-agent strings | Possible shared tooling | Low to medium |
| Network ranges | Hosting or geographic clustering | Medium |
| Referrer values | Often absent or forged | Low |
| Request frequency | Automation and scanning intensity | Medium to high |
The combination of synchronized timing and repeated URL order was more persuasive than any single field. A single IP making several suspicious requests may be an opportunistic scanner. Dozens of clients making the same requests in the same sequence indicate a different operational pattern.
This was also where Scratchy’s summaries were especially useful. Instead of manually searching every line, I could identify the high-volume paths, isolate time ranges, and focus detailed inspection on the intervals where activity intensified.
Following Infrastructure Clues
After identifying the behavioral pattern, I compared source addresses by network range and hosting provider. This did not establish attribution, but it helped show how the traffic was distributed. Some addresses were clustered in commercial hosting environments, while others appeared to come from residential or previously compromised networks.
I also checked whether the same addresses had interacted with other services on the server. Requests for static content, SSH-related activity, and scans against unrelated virtual hosts could reveal whether the web requests were part of a broader reconnaissance phase. The goal was to build a timeline, not to assume that every suspicious event had the same operator.
For larger investigations, I export Scratchy’s findings into additional scripts for deduplication and correlation. The same preference for focused, inspectable utilities appears in the Canyonero utility, where a narrow purpose makes behavior easier to understand and maintain. Small tools are often effective in forensic work because their output can be checked against the original evidence.
Testing the Attack Hypothesis
A hypothesis is only useful if it can survive attempts to disprove it. I compared the suspicious requests with normal traffic from search engines, monitoring services, and known application clients. Legitimate crawlers generally followed more predictable access patterns and respected site structure, while the suspicious group requested unrelated administrative and diagnostic paths.
I then checked response codes and response sizes. A coordinated scanner may generate many 404 responses while looking for a vulnerable file, but a successful 200 response can change the priority of the investigation. Redirects, forbidden responses, and unusually large responses also deserve attention because they may indicate authentication pages, defensive rules, or data exposure.
The findings supported a coordinated automated scan rather than a series of independent incidents. That distinction affected the response. Blocking individual addresses would provide short-term relief, but rate limiting, targeted deny rules, application updates, and monitoring for the same request sequence were more durable defenses.
Recommended Investigation Workflow
A repeatable workflow keeps log analysis from becoming a collection of guesses. My process for this investigation was:
- Preserve raw Apache logs and record the collection period before filtering.
- Use Scratchy to summarize clients, paths, status codes, timestamps, and request frequency.
- Group activity by behavior, including URL order and burst timing, rather than IP address alone.
- Compare suspicious traffic with known legitimate crawlers, monitors, and application clients.
- Apply defensive controls only after checking that they will not block required users or services.
The final step is documentation. I recorded the observed indicators, the reasoning behind the coordinated-attack assessment, and the controls applied afterward. This creates a baseline for future incidents and makes it easier to determine whether the same campaign returns with different addresses.
A useful report should also distinguish facts from interpretation. “Thirty clients requested the same four paths within two minutes” is an observation. “The clients were controlled by one attacker” is an interpretation that requires stronger evidence. Keeping those statements separate improves both technical accuracy and operational decision-making.
The value of Scratchy in this process was its ability to turn a large Apache log into an understandable set of relationships. It did not replace judgment, packet capture, system auditing, or application review. It made the first forensic question easier to answer: what behavior is repeating, and what does that repetition tell me?
Use the same method on your own server logs: preserve the evidence, summarize the activity, correlate behavior, and verify every assumption against the raw requests. With a focused analyzer such as Scratchy and a disciplined review process, scattered probes can become a clear timeline of coordinated activity—and a more informed security response.
