Building a Python script to harvest and analyse HTTP headers
HTTP headers travel with every request and response that crosses the web, carrying metadata that reveals how a server is configured, what software powers it, and how aggressively it protects visitors. For developers running audits across a portfolio of sites, harvesting these headers in bulk provides a snapshot of technical posture that manual spot checks cannot match. In Australia, where many government departments now align with the Australian Cyber Security Centre's Essential Eight framework and private firms benchmark against the same baseline, a script that pulls headers from dozens or hundreds of URLs becomes a practical compliance companion.
The approach below uses standard libraries and the popular requests package to walk through a plain-text list of addresses, record every header pair, and flag configurations that fall short of modern hardening expectations. It is the kind of lightweight tool that fits comfortably alongside the open-source utilities already catalogued on this site, and it can be extended as new header standards emerge.
Why HTTP headers matter for web auditing
Every response from a web server begins with a status line followed by a block of key-value pairs that tell the browser how to render content, how long to cache it, and which origins are permitted to embed the page. Security-focused headers such as Content-Security-Policy, Strict-Transport-Security, X-Content-Type-Options, and Referrer-Policy form a defensive layer that mitigates cross-site scripting, clickjacking, and protocol downgrade attacks. When these are absent or misconfigured, the consequences range from annoying browser warnings to genuine exploitation risk.
Australian site owners have additional motivation to audit headers because the Notifiable Data Breaches scheme and ACSC guidance both reference transport-layer protections explicitly. A small business in Brisbane running an online storefront, a Sydney fintech operating under CPS 234 obligations, or a Perth council delivering rate notices online all face similar exposure when a strict-transport-security directive is missing from responses. Harvesting headers in bulk turns these one-off observations into a quantifiable, repeatable check.
Setting up the Python environment
A clean virtual environment keeps the harvesting script portable and avoids polluting system packages. Creating one with python -m venv headers_env and activating it before installing dependencies is a habit picked up at every PyCon Australia session and reinforced through the Sydney and Melbourne Python user groups. Inside the environment, only two third-party packages are required for the basic version of the tool: requests for HTTP transport and rich if a coloured terminal output is desired.
The script itself leans on built-in modules for file handling, JSON serialisation, and CSV writing, which means it runs on any Python 3.9+ interpreter without external dependencies beyond requests. For developers who already maintain Python tooling, such as the Kodos regular-expression debugger catalogued elsewhere on this site, the workflow will feel familiar. Storing the URL list in a UTF-8 encoded text file with one entry per line keeps the input format transparent and easy to edit in any editor, including those used by teams collaborating across Australian time zones from Adelaide to Darwin.
Crafting the harvesting script
The core of the script is a function that accepts a URL, performs a HEAD request with a sensible timeout, and returns a dictionary mapping header names to their values. A HEAD method is preferable when the goal is purely metadata collection, because it avoids downloading response bodies that could be megabytes in size across a large list. Where a server rejects HEAD requests, the function can fall back to a GET request with a short streamed read so that the body is discarded while the headers remain intact.
Error handling deserves attention from the outset. Australian links to government services sometimes return 403 or 503 during planned maintenance windows, and overseas targets may rate-limit aggressive clients. Wrapping each call in a try-except block that catches requests.exceptions.RequestException ensures that one misbehaving URL does not abort the entire run. Recording the exception message alongside the headers gives the auditor a clear trail to investigate later, rather than a blank row in the output spreadsheet.
Storing and organising the results
Persisting the harvested headers in a structured format opens the door to deeper analysis. Writing each row to a CSV file with the URL, status code, server string, and a flattened selection of common security headers produces an artefact that can be opened in Excel, shared with stakeholders, or fed into a dashboard. A complementary JSON file keyed by URL preserves the full header set for callers who want every detail, including non-standard vendor headers that often leak internal hostnames or build numbers.
For larger crawls, dropping the data into a local SQLite database unlocks SQL queries that would be cumbersome in spreadsheets. Aggregating counts of CSP violations across every URL on a corporate portfolio, or filtering for sites that still serve Server: Apache/2.2 on an exposed endpoint, becomes a one-line query. This is especially useful for agencies in Canberra that must produce regular attestation reports, since the database can be archived as evidence of the audit trail.
Analysing security headers
Once the headers are stored, the real value emerges from analysis. A small dictionary that maps each security header to its expected presence or recommended value lets the script mark every URL as compliant, non-compliant, or partially compliant. Content-Security-Policy is rarely absent entirely on Australian banking sites, but its directives vary widely; flagging URLs that lack default-src 'self' or allow unsafe-inline scripts highlights genuine gaps that automated scanners sometimes overlook.
The Australian Signals Directorate publishes hardening guidance that mirrors the open OWASP recommendations, and aligning the analysis with that document makes the output meaningful to local security teams. Reporting on Strict-Transport-Security max-age values, X-Frame-Options settings, and Permissions-Policy declarations in a single summary gives technical leads a prioritised list of remediations. Exporting this summary as a Markdown table or PDF appendix keeps it presentation-ready for board packs that circulate from Melbourne headquarters to regional offices.
Scaling with AsyncIO and visualising findings
A synchronous script is perfectly adequate for a few dozen URLs, but crawling an entire portfolio would take hours. Refactoring the core fetch function to use aiohttp and asyncio.gather lets the tool issue hundreds of concurrent requests while respecting polite delays and respecting robots.txt where appropriate. Setting a semaphore around the gather call prevents runaway parallelism that could be mistaken for a denial-of-service attack by upstream providers or by vigilant network operations teams.
Visualisation closes the loop. Plotting the distribution of missing security headers with matplotlib or plotly, or generating a heat map that groups findings by Australian state, turns raw rows into a conversation starter. Adding timestamps in AEST so that scheduled crawls remain comparable, and logging each run to a simple append-only file, gives the auditor confidence that the most recent snapshot reflects the production environment rather than a stale cached view.
