Crawler Bench

How It’s Measured

Each crawler records the same sites, and replays under a synchronized clock. This is what each number means, and known limits.

The Test Machines

  • Mac mini (M1/Tahoe), 8 GB of RAM, macOS 26.6.2
  • Windows 10 (Hyper-V), 15.8 GB of RAM, Windows 10, running as a virtual machine

Crawls run one at a time, in isolation. Before each run the crawler is reset, and settings are applied or defaults accepted.

The Test Sites

Local sites are served from the test machine itself. The network drops out and the crawler becomes the bottleneck. A remote site holds the same files on a mirror reached over the internet. Links to other hosts are removed, and each crawler crawls the same single-host site.

Crawl Time

The clock starts when the crawl is submitted, by Start button or command. It stops when the crawler reports it’s done: a GUI’s completed state, or a CLI’s exit. A GUI is checked a few times a second, so its finish can trail slightly.

Configuring before the start, such as for JavaScript rendering or threads, shows in yellow on the replay and isn’t counted. Rank is by crawl time. A crawler stopped by a free edition’s URL limit didn’t finish, so it isn’t ranked.

Pages and Assets

Counts are what each tool reported, or for mirroring tools like Wget and HTTrack, what it saved. Tools classify and de-duplicate URLs differently, so totals vary for the same site. Beam Us Up SEO Crawler fetches HTML pages only, so no assets are reported.

Memory

Memory is the footprint of the crawler and every helper process it starts, such as a Java runtime or a browser engine for JavaScript, summed at each sample. On macOS that’s physical footprint (Activity Monitor); on Windows, private working set (Task Manager). The resident view counts only what’s held in RAM. Peak is the highest sample through the finish.

A spike between samples can be missed. A crawl shorter than one sample takes its peak from the operating system’s record of the process.

CPU, Disk, and Network

CPU sums the crawler’s processes as a percent of one core, so four busy cores show 400%. Disk is the read and write traffic of the crawler’s processes. Network is the whole machine, both directions, including anything else running. On a local site, loopback traffic counts twice: sent and received.

The Recordings

Each crawler’s window is recorded alone, scaled to 800 × 600, and aligned so every crawl starts together on the replay clock. Nothing on a replay ran at the same time.

Setups run free versions unless tagged Paid. InterroBot (JS), InterroBot (Turbo JS), InterroBot (Turbo), InterroBot CLI (JS), and InterroBot CLI (Turbo JS) use paid features, tested with a trial license.

Limits

Each setup is one recorded run, not an average. A remote site’s times include the internet path to the server, which varies run to run. Background work on the test machine shows up in machine-wide figures.