Skip to content

Feat/2748 paginated seed browsing - #2989

Open
Jackson Severino da Rocha (jacksonsr451) wants to merge 7 commits into
microsoft:mainfrom
jacksonsr451:feat/2748-paginated-seed-browsing
Open

Jackson Severino da Rocha (jacksonsr451) wants to merge 7 commits into
microsoft:mainfrom
jacksonsr451:feat/2748-paginated-seed-browsing

Conversation

@jacksonsr451

Copy link
Copy Markdown

Description

Implements #2748 by adding typed, read-only APIs for browsing and retrieving logical seed examples from existing seed memory.

This builds on the dataset identity and count contract introduced in #2762 and follows the browsing semantics discussed with Roman Lutz (@romanlutz) in #2748.

What changed

  • Added paginated logical seed example browsing with opaque, filter-bound keyset cursors.
  • Added deterministic ordering by the earliest stored date_added of the complete logical example, descending, with logical-example ID as the tie-breaker.
  • Uses prompt_group_id as the logical-example identity, falling back to the persisted seed ID for ungrouped records.
  • Added database-bounded filtering for:
    • modality
    • harm category
    • seed type
    • stored text search
  • Applies OR semantics within each filter family and AND semantics between filter families.
  • Evaluates filters at the logical-example level while returning all members of a matching example.
  • Added case-insensitive whole-value harm-category matching.
  • Added literal case-insensitive text search with % and _ treated as literal characters.
  • Added compact list projections with bounded previews and full detail projections for individual logical examples.
  • Added logical-example totals using the same predicates as the paginated query.
  • Added typed list/detail response models and HTTP routes:
    • GET /api/datasets/{selection_key}/seeds
    • GET /api/datasets/{selection_key}/seeds/{example_id}
  • Preserves persisted seed IDs, nullable group IDs, roles, sequence values, hashes, provenance, metadata, objective fields, template parameters, and persisted template state.
  • Added nullable persistence for is_jinja_template so browsing can distinguish templates without inferring template state from parameters.
  • Added side-effect-free projections for browsing without reconstructing seeds or groups.
  • Added direct bounded detail lookup rather than scanning paginated results.
  • Added SQLite and Azure SQL implementations for the required browsing predicates.
  • Preserved the existing MemoryInterface.get_seeds() behavior.

Safety and side-effect behavior

Browsing does not:

  • render templates;
  • reconstruct legacy simulated conversations;
  • load referenced template files;
  • load media contents;
  • generate conversations;
  • call models or targets;
  • write to seed memory;
  • fetch provider datasets;
  • load provider metadata solely to validate browsing selections.

Named browsing selections are resolved from datasets already represented in seed memory. A provider-only dataset that has not been persisted to memory is therefore not treated as browseable by these endpoints; this avoids provider/file loading during a read-only browsing request.

List previews also avoid exposing absolute filesystem paths or credential-bearing URLs.

Fixes #2748

Tests and Documentation

Added and expanded contract coverage across the Memory, service, persistence, and HTTP layers.

Coverage includes:

  • bounded database pagination;
  • deterministic ordering;
  • tied timestamps;
  • cursor continuation;
  • invalid and filter-mismatched cursors;
  • complete grouped-example retrieval;
  • cross-member filter matching;
  • OR-within-filter / AND-between-filter semantics;
  • whole-value harm-category matching;
  • literal % and _ text searches;
  • named and unnamed dataset isolation;
  • unknown dataset handling;
  • preview truncation;
  • full-detail fidelity;
  • original seed/group identifiers;
  • nullable template state;
  • template parameters;
  • side-effect-free browsing;
  • provider/file-loading guards;
  • direct bounded detail lookup;
  • SQLite and Azure SQL query behavior;
  • migration behavior and historical nullable template values;
  • HTTP dependency isolation.

Final focused validation:

  • HTTP seed browsing contract: 40 passed
  • Dataset service tests: 28 passed
  • Memory browsing validation: 108 passed
  • Migration tests: 71 passed
  • Relevant backend regressions: 158 passed
  • Ruff: passed
  • git diff --check: passed

Azure SQL integration was also validated against a real Azure SQL database after the SQL/schema corrections:

  • Azure SQL seed browsing integration: 1 passed

The final HTTP/service changes did not modify the SQL query, schema, or migration behavior after that Azure SQL validation.

JupyText was not run because this change does not modify notebooks or JupyText-managed documentation/examples.

@jacksonsr451

Copy link
Copy Markdown
Author

@microsoft-github-policy-service agree

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

FEAT GUI: Add paginated seed browsing API

1 participant