Skip to content

Show every Ruby's speed side by side on the results site - #252

Merged
JuanVqz merged 6 commits into
mainfrom
feature/results-site-view-2
Oct 7, 2026
Merged

JuanVqz merged 6 commits into
mainfrom
feature/results-site-view-2

Conversation

@JuanVqz

@JuanVqz JuanVqz commented Oct 5, 2026 •

Copy link
Copy Markdown
Member

Summary

  • Adds an "Across Rubies" view to the results site, and the site now opens on it: for every benchmark, how fast each idiom is on each Ruby, from 2.1 to head, JRuby and TruffleRuby, with and without the JIT.
  • Every build of a benchmark runs on the same machine, one after the other, so i/s can be compared across Rubies. Today each Ruby runs as its own CI job on its own machine (6 CPU models in one run), which only allows comparing idioms inside one Ruby.
  • Each workflow now does one thing. A new weekly workflow, Results site, does those runs in 6 shards and publishes the whole site, both views, from them. benchmarks.yml only checks that every benchmark runs on every Ruby, and no longer builds or publishes anything.

@JuanVqz
JuanVqz force-pushed the feature/results-site-view-2 branch 4 times, most recently from 2013d53 to ae14343 Compare October 6, 2026 19:43
The results site can only compare idioms inside one Ruby today: each Ruby
runs as its own CI job on its own machine, so i/s across Rubies is mostly
the machine. script/run_cross_ruby.rb runs every build of a shard's files
on one machine, one after the other, so their i/s can be compared.

The builds come from the rake matrix in benchmarks.yml, so a Ruby added to
CI is measured here too. The newest released MRI is the reference: it runs
first, again after every 3 builds, and last. It is the same code every
time, so its passes show how steady the machine was during the run.

A Ruby's interpreter and JIT builds run back to back, so each image is
fetched once and removed after its last pass (with --fresh-images, in CI),
which keeps the runner's disk free.

The collector records the shard, the pass, whether it is the reference and
when each report finished; those fields are empty for the regular CI jobs.
A new workflow, Results site (results-site.yml), runs
script/run_cross_ruby.rb in 6 shards, each on its own runner: every 6th
benchmark file, all 26 builds. A shard takes about 3 hours, under the 6
hour job limit, so it runs once a week (Sunday 03:00 UTC) and by hand.
Each shard uploads its results.
The site builder now reads a run of script/run_cross_ruby.rb and builds
both views from it, into across-rubies.json and does-it-hold.json
(named after the views). across-rubies.json has, per benchmark, every build's i/s
and error as measured and relative to the reference's median, and how much
the reference varied over its passes (standard deviation over mean). A
benchmark over 5% is marked noisy; when the typical benchmark is, the page
warns that small differences between Rubies may be the machine.

"Does it hold?" comes from the same run: every build's result, and the
reference's middle pass. Each verdict still compares reports inside one
run, and now every build of a benchmark shares one machine.

On the first run on GitHub runners the reference varied 2.3% on the typical
benchmark, with no drift over the run (last third within 0.1% of the
first). Correcting each build by the reference's speed at the time made
the lines jumpier (7.6% against 7.0% between MRI versions), as it did in
two local pilots, so the numbers are shown as measured.
The page gets a second view, Across Rubies, and opens on it. Per
benchmark: a chart with the Rubies on the x axis (MRI as a line, JRuby
and TruffleRuby as points), one line per report with the claimed one in
green, a JIT switch with the interpreter dashed under it, i/s or relative
to the reference, a table of every build, and how steady the machine was.

Two Rubies count as same-ish when their error bars overlap or their gap is
within how much the reference varied on that machine: runs minutes apart
vary more than benchmark-ips' error bars say. "Does it hold?" now says
every build of a benchmark ran on one machine, so its columns compare.

Both views keep their choices in the URL (view, and build, which they
share) and leave each other's alone, also on load.
Results site now builds the site from its six shards and deploys it, only
for a complete run on main. Before deploying it checks the last successful
run: if that run is newer (a re-run of an old run, or two runs finishing
out of order), it does not publish over its results.

benchmarks.yml goes back to checking that every benchmark runs on every
Ruby: no site preview and no result files, which nothing reads any more.
Each workflow now does one thing, and the site has one source.

A merged benchmark reaches the site with the next weekly run, or sooner by
running Results site by hand.
CONTRIBUTING: how to run a few benchmarks on a few builds and build the
site from them, which builds to leave out on an Apple silicon Mac, and
what each workflow does.
@JuanVqz
JuanVqz force-pushed the feature/results-site-view-2 branch from ae14343 to 491f640 Compare October 7, 2026 02:49
@JuanVqz
JuanVqz marked this pull request as ready for review October 7, 2026 04:28
@JuanVqz
JuanVqz merged commit 06c3bc9 into main Oct 7, 2026
29 checks passed
@JuanVqz
JuanVqz deleted the feature/results-site-view-2 branch October 7, 2026 04:28
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant