Benchmark in Milliseconds

(matklad.github.io)

36 points | by surprisetalk 22 hours ago

6 comments

  • vlovich123 1 hour ago
    Really depends on what you benchmark and how reliable you want the measurement and what domain you are benchmarking. For example, criterion can sometimes spend quite a bit of time because it needs to stabilize the measurements. The author’s claim is “well you don’t need that accuracy” but I’ve seen wasted time chasing ghosts or making claims on performance improvements that were either neutral or net negative due to this noise.

    > Anything faster than, say, 10ms risks being skewed by fixed costs (e.g, interpreter startup).

    Sounds like the author’s experience is strictly in Python. For example with Java you have to make sure the JIT has sufficient optimized your program.

    Additionally there’s plenty of situations where it can take a really long time to generate a representative dataset worth benchmarking and it can take time to evaluate the performance (eg databases). Short and quick microbenchmarks can be useful as building points, but at some point you need to evaluate steady state performance of the full thing. Other domains this comes up with is game rendering performance where a 300ms sample tells you nothing about whether you have frame drops after minute 25 or have a memory leak.

  • vardump 53 minutes ago
    Benchmarking like that is often broken because of continuous CPU core clock speed adjustments, system interrupts, SMIs, etc.

    I tried to fix it by switching hyperthreading off, playing with the scaling governor, boost, setting a CPU frequency to no avail. The jitter was too much and the results were not reproducible, so I just gave up.

    Of course your mileage may vary; this was on an AMD Zen 3 CPU.

    • hliyan 7 minutes ago
      Used to do this sort of thing for computations that needed to run in the 10 microsecond range (HFT stuff), circa 2008. Had very predictable results because:

      a) language was not garbage collected (C++)

      b) we avoided heap lock contentions in critical paths by pre-allocating object pools at startup

      c) I/O operations were offloaded to separate threads, connected by mutex locked linked lists

      d) processing thread was bound to its own CPU core

      That's about as deterministic as we could get.

  • veritron 35 minutes ago
    if you are doing benchmarks using interpreted languages but care about smaller timescales perhaps your problem is using interpreted languages.
  • winwang 58 minutes ago
    Love it. Unfortunately, for some benchmarks, it can be bit difficult to get representative inputs which take hundreds of milliseconds. I'm curious as to why the author doesn't loop 1-10ms inputs to deal with variance? Which also deals with startup costs. Rust microbench harnesses were already good at this stuff-out of-the-box (run-to-run variance, etc) several years ago.
  • hyperpape 12 minutes ago
    This has to be read in terms of https://xkcd.com/2400/. If you're chasing low-hanging fruit, and searching for a 3x or 10x speedup on something that has never been optimized, this advice probably can be ok.

    The harder you push, and the more you need to start finding smaller improvements, the more this advice becomes a rule of thumb you can't rely on.

  • jsd1982 57 minutes ago
    Assuming the subject of the benchmark is a web/API request here, otherwise the advice does not really apply. Milliseconds would be too large for benchmarking GPU- or CPU-intensive work.