How It’s Measured
Each crawler records the same sites, and replays under a synchronized clock. This is what each number means, and known limits.
The Test Machines
- Mac mini (M1/Tahoe), 8 GB of RAM, macOS 26.6.2
- Windows 10 (Hyper-V), 15.8 GB of RAM, Windows 10, running as a virtual machine
Crawls run one at a time, in isolation. Before each run the crawler is reset, and settings are applied or defaults accepted.
The Test Sites
- Rust Book, Local Server: 114 HTML pages and 176 files at rust-book-local.crawlerbench.com
- Rust Book, Remote Server: 114 HTML pages and 176 files at rust-book-remote.crawlerbench.com
- GIMP 3.0 Manual, Local Server: 711 HTML pages and 2,616 files at gimp-local.crawlerbench.com
- GIMP 3.0 Manual, Remote Server: 711 HTML pages and 2,616 files at gimp-remote.crawlerbench.com
Local sites are served from the test machine itself. The network drops out and the crawler becomes the bottleneck. A remote site holds the same files on a mirror reached over the internet. Links to other hosts are removed, and each crawler crawls the same single-host site.
Crawl Time
The clock starts when the crawl is submitted, by Start button or command. It stops when the crawler reports it’s done: a GUI’s completed state, or a CLI’s exit. A GUI is checked a few times a second, so its finish can trail slightly.
Configuring before the start, such as for JavaScript rendering or threads, shows in yellow on the replay and isn’t counted. Rank is by crawl time. A crawler stopped by a free edition’s URL limit didn’t finish, so it isn’t ranked.
Pages and Assets
Counts are what each tool reported, or for mirroring tools like Wget and HTTrack, what it saved. Tools classify and de-duplicate URLs differently, so totals vary for the same site. Beam Us Up SEO Crawler fetches HTML pages only, so no assets are reported.
Memory
Memory is the footprint of the crawler and every helper process it starts, such as a Java runtime or a browser engine for JavaScript, summed at each sample. On macOS that’s physical footprint (Activity Monitor); on Windows, private working set (Task Manager). The resident view counts only what’s held in RAM. Peak is the highest sample through the finish.
A spike between samples can be missed. A crawl shorter than one sample takes its peak from the operating system’s record of the process.
CPU, Disk, and Network
CPU sums the crawler’s processes as a percent of one core, so four busy cores show 400%. Disk is the read and write traffic of the crawler’s processes. Network is the whole machine, both directions, including anything else running. On a local site, loopback traffic counts twice: sent and received.
The Recordings
Each crawler’s window is recorded alone, scaled to 800 × 600, and aligned so every crawl starts together on the replay clock. Nothing on a replay ran at the same time.
Free and Paid
Setups run free versions unless tagged Paid. InterroBot (JS), InterroBot (Turbo JS), InterroBot (Turbo), InterroBot CLI (JS), and InterroBot CLI (Turbo JS) use paid features, tested with a trial license.
Limits
Each setup is one recorded run, not an average. A remote site’s times include the internet path to the server, which varies run to run. Background work on the test machine shows up in machine-wide figures.