Skip to content

Stream archives when listing and reading browsed files - #408

Open
abhinavgautam01 wants to merge 1 commit into
git-pkgs:mainfrom
abhinavgautam01:perf/streaming-archive-browse-386
Open

abhinavgautam01 wants to merge 1 commit into
git-pkgs:mainfrom
abhinavgautam01:perf/streaming-archive-browse-386

Conversation

@abhinavgautam01

Copy link
Copy Markdown
Contributor

Closes #386

Problem

Browsing a cached artifact buffered the whole archive in openArchive. The TAR readers also kept every expanded file body in memory. For non-npm packages the archive was parsed twice: once to detect a common root directory, then again with that prefix stripped. So listing one directory or viewing one small file could allocate memory for the entire expanded archive.

Changes

Streaming reader (internal/server/browse_archive.go)

Directory listings and file reads now use archives.OpenStream:

  • Listings keep only entry metadata. Bodies are never held in memory.
  • File reads stop at the first matching entry and copy it straight to the response.
  • npm keeps its package/ prefix stripping and needs a single pass.
  • Other ecosystems keep automatic common-root detection. A first pass reads entry names, stopping early as soon as two different top-level names appear. A file read then reopens storage for the second pass.
  • ZIP and conda need random access in the archives library. Their input is buffered once, within the existing input-size cap and shared by both passes so it is never read twice.

The listing logic mirrors archives.Reader.ListDir combined with the prefix wrapper, so paths, synthesized directory entries, ordering and duplicate handling all stay the same.

Limits

The existing 512 MB input-size cap stays. These limits are now explicit:

Limit Value
Expanded bytes across all entries 512 MB
Bytes per entry 512 MB
Entry count 100,000

They match what the buffered reader enforced implicitly, so no archive that browsed before is rejected now. A limit failure returns 500 with archive exceeds browse limits and is logged at warn level. Other archive errors still return failed to open archive.

Unchanged

  • Supported formats, path handling, Content-Type sniffing, Content-Security-Policy, X-Content-Type-Options and Content-Disposition headers.
  • Missing files still return 404.
  • The version diff still uses the buffered openArchive, because diff.Compare needs a random-access archives.Reader.
  • No API or Swagger changes.

Results

Fixture: a gzip tarball with 1,025 entries and 32 MiB expanded (92 KB compressed).

Operation Before After
List root directory 86.1 MB, 36 ms 0.88 MB, 17 ms
Read one small file 86.1 MB, 32 ms 0.62 MB, 12 ms

From BenchmarkBrowseLargeArchive (B/op and ns/op).

Tests

New tests in internal/server/browse_archive_test.go all go through the public browse endpoints:

  • Listings and prefix stripping for tar.gz with a root directory, extensionless tar.gz, npm, flat tar.gz, ZIP with a root directory, extensionless ZIP, flat ZIP and gem.
  • Parity with the old buffered reader for nine directory path variants ("", /, src, src/, /src, nested, missing, file path).
  • File contents and headers, 404 for missing files and that the first of two duplicate entries wins.
  • Limit failures for entry count, entry size, expanded size and input size, on both the streamed TAR path and the buffered ZIP path.
  • Memory: the large fixture is listed and one small file is read. The test asserts each streamed request allocates under a quarter of the expanded size, while the buffered open allocates more than the whole of it.

gofmt, go vet, golangci-lint (v2.13.1), go test -race ./... and Swagger regeneration all pass locally.

Notes

  • Because a file read stops at the first match, a corrupt or over-limit entry after the requested file is no longer detected for that request. The buffered reader rejected the whole archive in that case. This is inherent to streaming and listings still read every entry.
  • Conda has no fixture. Building one needs zstd, which would make klauspost/compress a direct dependency in go.mod. ZIP covers the same buffered path.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Use streaming archive readers for browsing and file extraction

1 participant