This repository contains the Hugo source for the EleutherAI website and blog.
- Main site: www.eleuther.ai
- Blog: blog.eleuther.ai
- Publication source of truth: EleutherAI papers Google Sheet
The main site and blog share templates, styles, data, and assets, but Hugo builds them as separate sites for separate Netlify deployments.
Deployment: Any changes made to main will be automatically deployed via Netlify. This can take a few minutes at times, but changes should be reflected in the live website pretty fast
The build requires Python 3.12 or newer, Make, and Hugo Extended 0.158.0. The supported versions are recorded in .python-version and .hugo-version.
For a new local checkout, install the development dependencies in a virtual environment:
python3 -m venv .venv
source .venv/bin/activate
python3 -m pip install -r requirements-dev.txt
playwright install chromium
make check-toolsPlaywright is used only for refreshing the Hugging Face download metric. Offline builds and tests do not launch a browser.
From the repository root, build the main site with current online data:
make buildThe static output is written to public/. To preview the complete website locally, start the main site and blog together:
make serve-all-offline
open http://127.0.0.1:8067/The main site runs on port 8067 and the blog runs on port 8068. Navigation automatically moves between the two local previews instead of opening the production websites.
For normal editing with Hugo's live-reloading server:
make serve-allThis refreshes the online data before starting both servers.
Build both the main site and blog:
make build-allThe individual preview commands remain available when only one deployment is needed:
make serve-offline PORT=8067
make serve-blog-offline PORT=8068When internet access is unavailable, use the checked-in data snapshots:
make build-offline
make build-all-offlineOffline builds do not check the publication Sheet, refresh online metrics, or query YouTube. They validate and use the checked-in snapshots instead.
Most text should be added to Markdown under content/ or structured YAML under data/. The files under layouts/ determine how that content is arranged. CSS and public assets live under site-page.css, static/, and assets/.
Files directly under data/ and data/projects/ are maintained by people. Everything under data/generated/ is written by the build scripts and must not be edited directly. Change the relevant Google Sheet, Markdown, YAML, external source, or generator and rebuild.
| Content | Edit here | Notes |
|---|---|---|
| Homepage hero and section structure | layouts/index.html |
The homepage's main prose and arrangement currently live in the template. |
| Homepage metrics, Current Research links, manual news | data/home.yaml |
Generated metric values replace entries that have a source key. Recent outputs are generated from the papers Sheet and blog front matter. |
| About page | content/about.md |
Donor logos and links are stored separately in data/donors.yaml. |
| Community page | data/community.yaml |
content/community.md contains only the page front matter. |
| Community reading-group cards | data/reading_groups.yaml |
One entry per reading-group series with its YouTube playlist ID (from the channel's playlists page). The build refreshes data/generated/community_reading_groups.json with each series' newest recording; use the repository's refresh-community-videos skill when maintaining or troubleshooting it. |
| Staff page | data/staff.yaml |
teams mirrors the org chart; give each person a team key. Store portraits locally under static/assets/staff/. |
| Research page and research areas | data/research/areas.yaml |
Area prose, key projects, current directions, and the Major Projects list. Each area also needs a stub under content/research/. The intro paragraphs are in content/research/_index.md. Paper lists per area are generated from the Sheet's Area column via research_area_filters.csv. |
| Research Library | Google Sheet | content/papers.md only defines the route and layout. Do not hand-edit the rendered paper list. |
| SOAR page | data/soar.yaml |
The page title and description are in content/soar.md. |
| News post | content/news/ |
Dated news posts can enter the homepage Latest feed. |
| Manual homepage news item | data/home.yaml under manual_news |
Give the item a title, URL, and date if it should participate in chronological sorting. |
| Blog post | content-blog/ |
Add Markdown front matter with at least title and date. Write math as plain TeX inside $…$ or $$…$$; MathJax loads automatically on pages that contain it. |
| General standalone page | content/ |
Most ordinary Markdown pages use a template under layouts/_default/. |
| Header and navigation | layouts/partials/header.html |
Shared by the main site and blog. |
| Footer | layouts/partials/footer.html |
Shared by the main site and blog. |
| Shared styling | site-page.css |
Page-specific styles are under static/, such as research-library.css. |
| Logos and source images | assets/ |
Hugo mounts selected assets into the built sites. |
| Directly served assets | static/ |
Files retain their paths in the generated site. |
public/andpublic-blog/are generated build outputs. Do not edit them./blog/andblog.htmlremain under review as possible deployment artifacts. Do not change or remove them until that deployment is confirmed.generate_blog_pages.pybelongs to the older standalone blog export and is not called by the current Makefile.
The Makefile coordinates data refreshes and Hugo. The normal main-site build is:
make build
-> python3 generate_hugo_data.py
-> python3 scripts/refresh_youtube_reading_groups.py
-> hugo --cleanDestinationDir
Each stage has a different purpose.
scripts/refresh_youtube_reading_groups.py reads data/reading_groups.yaml, which lists each reading-group series and its YouTube playlist ID. For every series it fetches the playlist feed (or, with YOUTUBE_API_KEY set, the full playlist through the Data API), takes the newest recording, and writes the series name, a link into the playlist at that recording, the session title, date, and a one-sentence overview from the official description to data/generated/community_reading_groups.json. A series with no playlist_id yet can set title_pattern to pick the newest matching channel upload instead; the freshness report flags this as live_limited until the playlist ID is added.
The checked-in JSON file is the deterministic offline fallback. A live build warns and uses that snapshot if YouTube is temporarily unavailable; it fails only when neither online discovery nor a valid three-video snapshot is available. make build-offline never queries YouTube and validates the snapshot before Hugo runs. make verify-live requires YOUTUBE_API_KEY so it can prove that the selection came from the complete uploads history.
The normal build runs this script online. It fetches external data, normalizes publication records, derives display fields, and writes JSON for Hugo.
The script downloads a cache-busted CSV export of the papers tab and writes it to:
eleutherai_papers.csv
The script verifies the full expected header set before replacing the local snapshot. If the Sheet cannot be downloaded, or its schema changes unexpectedly, a live build stops instead of silently claiming that stale publication data is current.
The principal Sheet columns currently used are:
| Sheet column | How the website uses it |
|---|---|
Title |
Display title and record identity |
Date |
Determines whether a paper enters the public library; supplies its displayed year and chronological sort |
Highest Impact |
Selects papers for the generated homepage Recent outputs candidate set; checked rows export as TRUE |
Display Authors |
Author line shown in the Research Library |
All Authors |
Complete semicolon-separated author list used for author search and metadata generation |
Area |
Semicolon-separated research-area metadata used by generated collections and filters |
Paper Link |
Makes the paper entry clickable |
Conference or Journal |
Primary archival venue; blank values fall back to a workshop or arXiv |
Workshop |
Workshop appearance; may coexist with an archival venue |
Superlative |
Oral, spotlight, best-paper, and runner-up text and symbols |
A row must have both a normalized title and a parseable Date to appear in the main Research Library and publication count.
The script cleans OpenReview tracking parameters and uses a small URL override table for known papers whose links were missing or incorrect in earlier data.
Venue logic is derived from Conference or Journal and Workshop:
- Papers use the conference or journal, then a workshop if no archival venue exists, then arXiv as a fallback.
- Papers with both an archival venue and a distinct workshop can produce separate appearances in grouped publication data.
The generator maps the simplified Sheet schema into its internal publication model:
Datecontrols inclusion and the one-record-per-paper library chronology.- The same date is used for conference and workshop appearance sorting because the Sheet intentionally maintains one date per paper.
- arXiv groups are normalized to January 1 so they are treated as the earliest chronological point in their year and appear after later conference and workshop groups in the newest-first display.
- Workshop sort dates receive a seven-day offset so a corresponding main conference is presented above its workshops.
That final ordering rule is editorial rather than historical: it keeps main conferences visually prominent even when their workshops occurred later.
The script cleans the Superlative field and associates an award with the appropriate conference or workshop appearance. Commas inside a single award description are preserved. For the Research Library it chooses one marker using this precedence:
- runner-up or finalist
- best paper
- spotlight
- oral
The generated marker value is rendered by layouts/partials/paper-marker.html, which keeps symbols and accessible labels consistent across the site.
The website needs two related but different publication representations.
all_papers() can create more than one record for a paper when it appeared at both a workshop and an archival venue. Each appearance contains:
- title and URL
- publication year
- venue and venue year
- venue kind: conference, workshop, or arXiv
- normalized conference family
- group sort date
- appearance-specific distinctions
These records feed the chronological venue-grouped publication components.
library_papers() creates exactly one record per dated paper. In addition to title, URL, year, and venue, each record contains:
- authors, displayed as last names and shortened after four names
- status
- primary and additional areas
- lead organization
- EleutherAI contact
- artifact type, currently
Paper - distinction marker and text
The current public Research Library does not display every stored metadata field, but the fields remain available to its filters and templates.
grouped_papers() groups publication appearances first by conference family and year, then by track or workshop. Examples include a main ICML conference group, a named workshop under ICML, and an arXiv preprint group.
Within a year, the generated order favors:
- main conferences
- workshops
- arXiv
Papers inside each venue are sorted in reverse chronological order.
The script creates data/generated/research/area_papers.json in two ways:
- Training Dynamics uses the hand-curated titles, summaries, and optional venue labels in
research_area_papers.csv. - Other areas use the broad-area, include-term, and exclude-term rules in
research_area_filters.csv.
The keyword filters search the paper title, superlative, conference or journal, workshop, and semicolon-separated Area values. These generated sets are useful data, but individual research-area pages are intentionally not exposed in the current public release.
The datasets key in this file is a filtered set of papers about datasets. It is not a catalog of released dataset artifacts, and the build does not currently query Hugging Face for a dataset inventory.
The generator writes data/generated/home_generated_metrics.json from three sources.
It counts unique titled Sheet rows with a valid Date, rounds the result down to the nearest 25, and appends +. The resulting object uses the key publication_count.
It fetches EleutherAI's Google Scholar profile, extracts the total citation count, rounds down to the nearest thousand, and appends + when the homepage renders it. The successful response is cached in data/generated/home_scholar_metrics.json.
If Scholar cannot be refreshed, the script uses the cached value. If no cache exists, it omits the metric rather than failing the full build.
It obtains the Hugging Face model-download value in this order:
HF_MODEL_DOWNLOADSenvironment variable, when set- the rendered EleutherAI publisher analytics page through Playwright
- the checked-in cache in
data/generated/hf_model_downloads_cache.json
If neither a live value nor a cache is available, the metric is omitted. This is a model-download metric only; it does not count dataset downloads.
The homepage reads data/home.yaml. Metric entries with a source such as publication_count, scholar_citations, or model_downloads are filled from data/generated/home_generated_metrics.json during Hugo rendering.
The generator scans content-blog/ recursively. It skips the section index and posts marked draft: true, then reads each post's title and date. It writes a sorted list to data/generated/blog_posts.json with production blog-subdomain URLs.
The homepage combines these blog entries with content/news/ and dated manual_news entries from data/home.yaml. layouts/index.html sorts the combined feed and displays the four newest items.
The first pass writes:
| Generated file | Contents |
|---|---|
data/generated/research/papers.json |
All publication appearances, including separate workshop and conference appearances |
data/generated/research/homepage_papers.json |
The first 10 publication appearances |
data/generated/research/paper_groups.json |
All appearances grouped by venue family, year, and track |
data/generated/research/homepage_paper_groups.json |
Grouped data from the first 30 appearances |
data/generated/research/area_papers.json |
Curated and keyword-filtered research-area paper sets |
data/generated/research/library_papers.json |
One searchable record per dated paper |
data/generated/research/paper_collections.json |
Named publication collections such as SOAR |
data/generated/home_generated_metrics.json |
Publication, citation, and model-download homepage values |
data/generated/home_scholar_metrics.json |
Last successful Scholar response used as a fallback |
data/generated/hf_model_downloads_cache.json |
Last supplied or successfully fetched download value |
data/generated/blog_posts.json |
Dated blog entries used by the homepage Latest feed |
data/generated/home_recent_outputs.json |
Four newest items from highest-impact papers and published blog posts |
data/generated/community_reading_groups.json |
One card per configured reading-group series, newest recording first |
Do not edit these files directly. Edit their source data and rebuild.
The Sheet's semicolon-separated all authors column is the sole source for complete publication authorship. generate_hugo_data.py puts every full name into the hidden Research Library search index while continuing to display the shorter Display Authors line. A live or offline build stops with a list of affected titles if any paper has a blank all authors cell; fix the Sheet rather than adding a local author override.
After data generation, hugo --cleanDestinationDir renders the main site with hugo.toml and writes public/.
Important data connections include:
layouts/index.htmlreadsdata/home.yamland generated homepage data underdata/generated/.layouts/_default/research-library.htmlreadsdata/generated/research/library_papers.json.layouts/_default/research.htmlandlayouts/partials/publication-groups.htmlread the grouped publication data.layouts/_default/community.htmlreadsdata/community.yaml.layouts/_default/staff.htmlreadsdata/staff.yaml.layouts/_default/soar.htmlreadsdata/soar.yamland the SOAR paper collection.
make build-all performs the complete main-site process above and then builds the blog with hugo-blog.toml into public-blog/. A standalone make build-blog does not refresh publication data first; it uses the generated data already present in the checkout.
For ordinary content that does not depend on external data:
make verifyFor publications, homepage metrics, news aggregation, or other externally sourced data:
make verify-livemake verify runs the tests, regenerates checked-in snapshots offline, builds both deployments with Hugo warnings treated as errors, checks local links and assets, and fails if generated snapshots differ from HEAD or the diff contains whitespace errors. GitHub Actions runs the same command.
make verify-live additionally requires every external source to be current. Set YOUTUBE_API_KEY for complete YouTube uploads discovery and either install Playwright or provide HF_MODEL_DOWNLOADS. Run make freshness after any build to see which sources were fetched live and which used snapshots.
Review the affected main-site and blog pages locally before committing generated changes.