Using Scratchy to Detect Crawlers and Bots in Apache Logs
Apache access logs contain a detailed record of how visitors and automated clients interact with a website. Search engines, monitoring services, feed readers, vulnerability scanners, scraping scripts, and malicious bots all leave traces in the same stream of requests. Separating these behaviors is essential for understanding traffic and protecting server resources.
Scratchy, the Apache log analyzer created by Phil Schwartz, provides a practical way to examine those traces. Rather than relying on a single user-agent string, you can use its reports to compare request volume, requested paths, response codes, referrers, and access patterns. This turns a large log file into evidence that can be reviewed and acted upon.
The tool is especially useful for periodic analysis. Run it against a current or rotated Apache log, identify unusual clients, and then compare the results with later reports. Over time, this creates a clearer picture of normal visitors, legitimate crawlers, and unwanted automation.
What Scratchy reveals in access logs
A crawler investigation begins with the basic fields recorded by Apache: client address, timestamp, request method, URL, protocol, status code, response size, referrer, and user-agent. Scratchy can organize this information into useful summaries, making it easier to spot clients that request hundreds of pages while producing little meaningful referral traffic.
A high request count does not automatically indicate abuse. A search engine may crawl many pages while returning a recognizable identity and following a consistent path. A bot that requests missing administrative files, exposed configuration names, or random query strings presents a different risk. Scratchy helps expose these differences by showing what each client requested and how the server responded.
Prepare Apache data for reliable analysis
Use complete access logs whenever possible, including the standard or combined Apache format. A short sample may hide periodic crawlers, while a single hour can make a normal indexing service look aggressive. Reviewing a full day or week gives request frequency and timing enough context to support better decisions.
Log rotation should be considered before running a report. Preserve the original files, work from a copy, and record the date range and virtual host represented by each file. If several websites share a server, separate their logs before analysis; otherwise, a busy domain can obscure the behavior of a smaller site.
Privacy also matters. IP addresses and URLs may contain sensitive information, including query parameters. Restrict access to raw logs and remove unnecessary personal data before sharing reports. Scratchy is an analysis utility, not a replacement for sound log-retention and access-control practices.
Distinguish search crawlers from suspicious automation
The user-agent field is a useful starting point, but it is easy to forge. A client can call itself Googlebot, Bingbot, or a browser while behaving like a scraper. Treat user-agent names as labels that require verification rather than proof of identity.
Compare the claimed identity with behavior. A reputable search crawler generally requests public pages, respects sensible crawl pacing, and returns patterns consistent with indexing. Suspicious automation may request login endpoints, backup files, framework probes, or URLs that do not appear in the site’s links. Repeated four-hundred responses, rapid sequential requests, and many unrelated paths are strong signals for investigation.
IP address data can add context, but it should be interpreted carefully. Hosting providers, proxies, and changing network ranges make simple address-based assumptions unreliable. When a crawler claims to belong to a known service, use that organization’s published verification guidance rather than blocking an entire address range based only on a name in the log.
| Signal | Likely interpretation | Follow-up |
|---|---|---|
| Recognizable user-agent with steady page requests | Possible legitimate crawler | Verify identity and review crawl rate |
| Many requests for nonexistent files | Scanner or misconfigured client | Inspect paths and consider rate controls |
| Repeated login or admin requests | Credential or discovery activity | Check authentication logs and block carefully |
| High volume from one address | Aggressive bot, proxy, or shared network | Compare timing, URLs, and status codes |
| Requests for public assets only | Browser, cache warmer, or monitoring service | Check referrers and normal site behavior |
| Rapid requests with changing user-agents | Evasion or scraping automation | Group by address, timing, and request pattern |
Read status codes as behavioral evidence
Status codes help explain what a bot was trying to accomplish. A crawler that mostly receives 200 responses may be reading public content. A client generating a large number of 404 responses could be following stale links, probing for known software files, or guessing paths. A concentration of 403 responses may indicate that access controls are already stopping unwanted requests.
Server errors deserve separate attention. A bot that repeatedly triggers 500-level responses may be exposing an application weakness, creating unnecessary load, or finding an input that the site does not handle safely. Scratchy’s summaries can help identify the relevant URLs and time periods, while Apache and application logs provide the deeper diagnostic detail.
Status codes should be combined with response sizes and timing when available. A client receiving small error pages at a high rate behaves differently from one downloading large public documents. Looking at several signals reduces the chance of blocking a legitimate service because of one unusual request.
Turn reports into a repeatable bot-audit workflow
Start with a baseline report during a period when the site is operating normally. Note the most frequent clients, popular paths, common user-agent strings, and typical proportions of successful and failed requests. This baseline makes future changes easier to recognize.
Next, isolate anomalies for manual review. Look for sudden increases in requests, unfamiliar user agents, repeated access to sensitive paths, and traffic that ignores the site’s normal content structure. Correlate the findings with deployment records, uptime monitors, search-console data, firewall events, and application authentication logs.
The process should end with a measured response. A verified crawler may need no action, while a noisy scraper could be slowed with server rules or application-level rate limiting. A scanner targeting login pages may justify a firewall rule or stronger authentication controls. Keep a record of the evidence and the action taken so that a later report can confirm whether the change helped.
Practical recommendations for better bot detection
Scratchy is most valuable when its output supports consistent review rather than one-time curiosity. Use the same date ranges and reporting categories where possible, then compare results across weeks or after a site change.
Apply these practices during each review:
- Analyze complete Apache access logs and record the time period covered.
- Group suspicious traffic by user agent, client address, requested path, and response code.
- Verify claimed search-engine identities before allowing broad crawler access.
- Treat high request volume as a signal for investigation, not automatic proof of abuse.
- Correlate access-log findings with firewall, application, and authentication events.
Avoid blocking every unfamiliar bot. New research tools, accessibility services, uptime monitors, and legitimate crawlers can resemble ordinary automation. A targeted response based on repeated behavior is safer than a broad rule that harms useful traffic.
Make log analysis part of server maintenance
Crawler activity changes as a website grows, its URLs change, and new software vulnerabilities become common. A report that was accurate last month may no longer describe the current traffic mix. Periodic Scratchy analysis gives administrators a lightweight way to detect those changes before they become a performance or security problem.
Use the results to improve robots.txt guidance, remove unnecessary public endpoints, tune rate limits, and investigate recurring error responses. Keep verified crawler information separate from assumptions, and preserve representative reports so that traffic trends remain understandable.
Download Scratchy, run it against a carefully selected Apache log, and use the resulting evidence to identify which automated clients deserve access, monitoring, throttling, or blocking.
