Callgrinder shows which functions cost the most in your C/C++ application. Give it a command and it runs Valgrind's callgrind profiler, then prints the hottest routines. An optional JSON config groups functions into categories you care about.
Use it locally with ./callgrinder or in GitHub Actions with gemc/Callgrinder@v1. Costs are estimated CPU
cycles (CEst), not elapsed seconds. To measure how runtime changes with thread count, use
ThreadScale.
- Linux, including a Linux VM or container on macOS or Windows.
- Node.js 24+, Git, and Valgrind (including
callgrind_annotate). - Your application compiled in debug mode, with debug symbols (
-g; CMake:-DCMAKE_BUILD_TYPE=Debug).
Callgrinder itself has no npm dependencies or build step.
Compile your software in debug mode and check that your executable runs normally with its input files. Then clone Callgrinder and pass it that same command:
git clone https://github.com/gemc/Callgrinder.git
cd Callgrinder
./callgrinder './build/bin/myapp input.dat' \
--working-directory /path/to/your/project \
--name my-workload --output-dir profile-resultsReplace /path/to/your/project with your application's directory and ./build/bin/myapp input.dat with your
executable and its arguments. The executable and input paths are relative to --working-directory. Start with
a small workload because profiling is much slower than a normal run.
The terminal prints the Top routines table. Category rows are empty until you supply a category config. Results are saved under the Callgrinder checkout:
profile-results/callgrind.out.my-workload: the raw profile for QCachegrind or KCachegrind.profile-results/profile-my-workload.json: the parsed results used to generate a report.
Save the summary as Markdown and export the category table as CSV:
./callgrinder report --input-dir profile-results --output-dir profile-report
cat profile-report/summary.mdThis creates summary.md, categories.csv, and callgrinder.json in profile-report/. The CSV contains
category rows once you add a config.
Always pass --output-dir when running from the checkout: the default callgrinder directory name conflicts
with the callgrinder launcher file. Choose a different directory or profile name to keep earlier runs.
The command may contain {events}, {name}, and {run} placeholders. Quote the whole command as one
argument, and quote inner arguments that contain spaces:
./callgrinder 'gemc card.yaml -n {events} -gsystem="[{name: det, factory: ascii}]"' \
--name det --events 100 --output-dir profile-resultsCallgrinder runs the command as valgrind --tool=callgrind … bash -c "exec <command>", so exec makes the
shell hand its process to your binary and callgrind profiles the binary directly rather than the shell.
If you would rather run valgrind yourself — for example to pass complex, space-containing arguments as a
shell array with no re-quoting — hand Callgrinder the resulting file and it only summarizes:
./callgrinder --from-callgrind profile-results/callgrind.out.my-workload \
--name my-workload --output-dir profile-resultsCategories are optional. Save your category definitions as ci/callgrinder.json and pass
--config ci/callgrinder.json. Config paths are relative to the directory where you invoke Callgrinder.
You can add categories to an existing profile without rerunning your application:
./callgrinder --from-callgrind profile-results/callgrind.out.my-workload --name my-workload \
--config /path/to/your/project/ci/callgrinder.json --output-dir profile-results
./callgrinder report --input-dir profile-results --output-dir profile-reportEach category is either fixed (match names one entry symbol) or discovered (discover captures a
class in group 1 and reports one row per class found — e.g. every plugin of a kind):
{
"cost": "CEst",
"top_routines": 10,
"categories": [
{ "label": "Track swimming", "match": "G4PropagatorInField::ComputeStep" },
{ "family": "Digitization", "discover": "([A-Za-z_]\\w*)::digitizeHit" },
{ "family": "Field", "discover": "(GField_[A-Za-z0-9_]*)::GetFieldValue" }
]
}match— a regex naming one entry function; its inclusive and self cost become one row.discover— a regex whose group 1 captures a class; every matching class becomes its own row, labelled"<family>: <class>". Add"method": "…"to control the displayed symbol whendiscoveruses alternation.
config accepts a file path or inline JSON. See GEMC3 and GEMC2
for application-specific examples.
The optional project_callers array defines which routines count as project callers when looking upstream
through runtime functions. Each regex matches a routine name, source path, or binary/library path separately:
{
"project_callers": ["^MyApp::", "libMyPhysics", "/work/my-project/"],
"categories": [
{ "label": "Input", "match": "MyApp::read" }
]
}Without this setting, Callgrinder treats named routines outside common C/C++ runtime libraries and namespaces as project routines. This inference includes application dependencies and is independent of any framework. Use explicit patterns for custom runtimes or to focus on selected libraries. An empty array disables project caller matches. Caller reporting and this setting are upcoming in the next release.
Use a Linux runner with Valgrind already available. This example uses a self-hosted runner and a CMake
application; CMake and your compiler must also be available. Save it as .github/workflows/profile.yml,
replace the build commands with your project's debug build, and set command to your executable and inputs:
name: Profile
on:
workflow_dispatch:
permissions:
contents: read
jobs:
profile:
runs-on: [self-hosted, linux]
steps:
- uses: actions/checkout@v6
- name: Build your application in debug mode
run: |
cmake -S . -B build -DCMAKE_BUILD_TYPE=Debug
cmake --build build --parallel
- uses: gemc/Callgrinder@v1
with:
command: ./build/bin/myapp input.dat
name: my-workload
output-dir: profile-results
- uses: gemc/Callgrinder@v1
with:
mode: report
input-dir: profile-results
output-dir: profile-report
- uses: actions/upload-artifact@v7
with:
name: profile
path: |
profile-results/
profile-report/Once the workflow is on your default branch, open Actions → Profile → Run workflow. Read the tables in
the run summary and download the profile artifact for the raw profile and report files. Prepare any input
files your application needs before the Callgrinder step. Add config: ci/callgrinder.json when you have
defined your categories. The Action supplies its own Node.js runtime.
- Two tables with matching columns — the category table groups the configured entry routines; the
top-routines table selects individual routines with the highest
Self %, then orders them by inclusive% of run, largest first. Both show category, symbols with project caller names in parentheses, inclusive and self Mcycles,% of run,Self %, totalCalls, and direct callers with counts. Matching columns and caller attribution are upcoming in the next release. - Costs — inclusive cost includes callees and overlaps, so neither table's
% of runmay be added. Self cost counts only work directly in the named routines. Category rows aggregate their matched entries; overlapping category patterns can repeat work. Individual self shares sum to at most 100%, apart from rounding. Listed and remaining self-cost shares appear below the top-routines table. - Calls and direct callers — counts sum recorded incoming calls, including recursion and object copies. Each direct caller is listed with its own count, largest first. Categories count each matched routine once, including calls between matched routines. Zero means no recorded incoming calls.
- Project callers in parentheses — for a routine such as
mallocorstrtod, follow each upstream branch past runtime functions to its first project caller. The routine cell includes those function names, for examplestrtod (MyApp::readValue), identifying the application or library code responsible for the observed calls. Up to three distinct names appear;+N moremarks additional names. Full signatures and shortest hop distances are available in JSON. This compact display is upcoming in the next release. Ownership uses runtime-name inference orproject_callers, independent of any framework. Recursive cycles are bounded. Aggregate profiles describe observed edges, so upstream callers have no inferred call counts. - Source and caller fields are also included in JSON results for both tables.
—means missing information or no matching caller. Regenerate existing partial JSON reports with--from-callgrindto populate these fields; the application does not need to be profiled again. The category CSV retains its cost columns. - The tables omit source/package lists. JSON retains source metadata, and unresolved addresses are labelled with their object when available.
- The function-table parsing fix shipped in v1.0.6. Regenerate older partial JSON reports with
--from-callgrind; the application does not need to be profiled again. - Cost is CEst (
Ir + 10·L1_misses + 100·LL_misses), matching qcachegrind's cycle estimation. The report ends with a short qcachegrind reading guide.
profile(default) — run one command under callgrind; writecallgrind.out.<name>, a partial JSON, and a Job-Summary section.report— merge partial JSON files frominput-dirintosummary.md,categories.csv, and an aggregated JSON.discover— turn a JSONbenchmarksarray into a job matrix (one job per profile) for fan-out.
The reusable workflow .github/workflows/callgrinder.yml wires discover → profile (matrix) → report.
npm run check # node --check on the entry points
npm test # node --testNo build step and no npm dependencies — keep it that way. See releases/ for per-version notes.