How Scratchy Categorizes Bot Traffic vs Human Traffic in Apache Logs
Apache access logs record every request in a compact, machine-readable format. That raw stream can reveal popular pages, broken links, hostile probes, crawler activity, and the behavior of people browsing a site. The difficulty is that a log file does not explicitly label each visitor as human or automated.
Scratchy approaches this problem as a practical log-analysis task. It examines request details such as client identity, requested resources, response codes, and access frequency, then organizes those observations into useful traffic categories. The result is a clearer view of how a site is being used without requiring a full analytics platform.
The distinction is based on evidence rather than certainty. A browser can imitate a crawler’s user-agent, while a well-behaved search bot can resemble an ordinary visitor. Scratchy therefore provides a meaningful operational classification, not a claim about the biological identity of every client behind an IP address.
How Apache records visitor behavior
A standard Apache access entry commonly includes the remote address, authenticated identity, timestamp, request line, status code, response size, referrer, and user-agent string. Scratchy uses these fields as the building blocks for its reports. Each line becomes a small event that can be grouped with related requests.
A single request is rarely enough to identify a traffic source. A client requesting a page, a stylesheet, an image, and a script within a short period looks more like a browser session than a client requesting thousands of unrelated paths. Patterns across many lines are therefore more informative than any one field.
This approach also makes the reports useful for troubleshooting. A sudden rise in 404 responses may indicate a broken crawler, a vulnerability scanner, or outdated links. A large volume of successful requests for static files may represent normal page loading, an aggressive download tool, or a cache that is not working as expected.
Signals that separate automated clients
The user-agent field is one of Scratchy’s clearest indicators. Names associated with search engines, monitoring services, feed readers, download utilities, and command-line clients can be grouped as bots or automated tools. Generic values such as an empty user-agent or an unusual scripting-library signature are also worth examining.
Request rate and repetition add another layer of evidence. Automated traffic often requests pages at regular intervals, walks through many URLs quickly, or repeats the same path with little variation. It may also target administrative files, configuration names, and common software weaknesses rather than following the visible links of a site.
Status codes help refine the picture. A client generating a high percentage of 404 or 403 responses is unlikely to be conducting ordinary browsing. Repeated 500 responses can point to an automated process exposing an application problem, while a crawler that receives mostly 200 responses may be indexing legitimate content.
Scratchy’s categorization is best understood as heuristic. It relies on recognizable patterns in Apache logs instead of claiming to perform advanced behavioral detection or validating every crawler through external DNS checks. That keeps the utility lightweight and makes its output understandable to developers working directly with text log files.
Signals that resemble human browsing
Human visitors usually generate a sequence of related requests. They may open an HTML page, fetch its supporting assets, move to another page, and occasionally return to a previous resource. Referrer values can help connect these actions, although privacy settings, bookmarks, browser extensions, and direct navigation often leave the field empty.
Traffic volume alone does not prove automation. A popular page can create thousands of legitimate requests, especially when many people arrive at once. Conversely, a carefully written bot can slow itself down and use a realistic user-agent. Scratchy’s value is in showing the surrounding context so an administrator can interpret ambiguous activity.
Browser traffic also tends to produce a mixture of resource types. HTML documents, images, CSS, JavaScript, and media files appear in combinations that reflect page structure. A command-line scanner may request only application endpoints or guessed filenames, creating a very different resource profile even when it identifies itself as a browser.
The distinction matters for performance analysis. Human-oriented requests reveal which pages attract attention and where navigation succeeds or fails. Bot-oriented requests reveal indexing behavior, vulnerability scans, uptime checks, and other background activity that can dominate bandwidth or obscure genuine usage statistics.
Reading the categories in a report
Scratchy’s output becomes more useful when categories are compared instead of read in isolation. A count of requests from bots may look alarming until it is separated into search indexing, monitoring, failed probes, and ordinary automated downloads. Similarly, a high human count may include repeated page refreshes or traffic from shared networks.
| Evidence in the log | More consistent with human traffic | More consistent with bot traffic |
|---|---|---|
| Request sequence | Related pages and supporting assets | Repeated or guessed paths |
| User-agent | Common browser signature | Crawler, script, scanner, or blank value |
| Timing | Irregular navigation intervals | Rapid, fixed, or highly repetitive requests |
| Referrer | Internal links or search results | Missing, suspicious, or unchanged referrer |
| Status codes | Mostly successful page loads | Many 403, 404, or probing-related errors |
| Resource mix | HTML combined with images and scripts | Narrow focus on endpoints or filenames |
These indicators should be combined rather than treated as fixed rules. For example, a search engine may request HTML without images, producing a narrow resource mix while still behaving legitimately. A browser that has cached assets may also generate fewer supporting-file requests than expected.
The categories are especially valuable when investigating a time range. Compare a quiet period with a traffic spike, inspect the most active client identities, and then review the requested paths. This helps distinguish a genuine audience increase from a crawler burst or an automated attack.
Where classification can mislead
User-agent strings are self-reported and easy to change. Blocking or labeling every client that claims to be a bot can exclude real users, while trusting every browser signature can allow scripts to pass unnoticed. IP addresses have similar limitations because multiple people may share one address through a proxy, carrier network, or office gateway.
Privacy technologies further reduce certainty. Browsers may omit referrers, rotate identifiers, or limit observable details. Reverse proxies and content delivery networks can also alter which address Apache records unless forwarding information is configured correctly. Scratchy can summarize the data present in the log, but it cannot recover details that were never recorded.
For that reason, administrators should treat a category as a lead for investigation. Correlating Scratchy’s results with firewall events, web-server timing, application logs, and resource usage can establish whether an unusual client is harmless, inefficient, or hostile. The report supplies focus; operational decisions require context.
A practical workflow for useful results
Start with a representative Apache log rather than a short fragment collected during an incident. A longer period shows recurring crawlers, normal visitor rhythms, weekly patterns, and changes in error rates. Preserve the original file, then analyze a copy so that filtering and compression do not remove evidence.
Use the report to prioritize investigation in a consistent order:
- Review the largest request sources and their user-agent strings.
- Compare successful requests with 403, 404, and 500 responses.
- Look for rapid repetition, sequential URL guessing, and unusual paths.
- Separate recognized search or monitoring bots from unknown automation.
- Check whether apparent human sessions contain coherent page and asset sequences.
The same workflow can support capacity planning and security work. Developers can identify popular resources and broken routes, while system administrators can find clients that deserve rate limiting or further review. Other small utilities in a software portfolio, including the Canyonero project, reflect the same practical preference for focused tools that expose useful technical evidence.
Scratchy is most effective when its categories remain connected to the underlying requests. Keep the original timestamps, client addresses, paths, status codes, and user-agent values available for verification. That balance between summarized results and inspectable raw data makes the analysis reproducible.
Use Scratchy to turn Apache’s dense request history into an evidence-based view of browsing, crawling, and probing activity. Run it against representative logs, study the patterns behind each category, and use the findings to improve monitoring, performance analysis, and server security.
