diff --git a/.agents/plugins/marketplace.json b/.agents/plugins/marketplace.json index 7f45bee..54e4991 100644 --- a/.agents/plugins/marketplace.json +++ b/.agents/plugins/marketplace.json @@ -15,6 +15,18 @@ "authentication": "ON_INSTALL" }, "category": "Productivity" + }, + { + "name": "fxtr", + "source": { + "source": "local", + "path": "./plugins/fxtr" + }, + "policy": { + "installation": "AVAILABLE", + "authentication": "ON_INSTALL" + }, + "category": "Productivity" } ] } diff --git a/README.md b/README.md index 080d18d..ec95c5b 100644 --- a/README.md +++ b/README.md @@ -7,6 +7,14 @@ A collection of plugins for Codex. | Plugin | Description | |--------|-------------| | [docent](./plugins/docent) | Docent AI analysis tools for Codex | +| [fxtr](./plugins/fxtr) | Skills for writing, running, and viewing fxtr experiments, and for calling models from them with behaviors | + +The fxtr plugin includes two skills, each with a copy of the documentation it links from +[docs.transluce.ai](https://docs.transluce.ai): +- **fxtr** - Setting up, writing, running, and viewing fxtr experiments +- **behaviors** - Calling language models from fxtr experiments with the behaviors library + +The fxtr plugin is generated from the fxtr repository by `pnpm sync:plugin`; don't edit it by hand. ## Installation @@ -16,8 +24,9 @@ Add this marketplace to Codex: codex plugin marketplace add TransluceAI/codex-plugins ``` -Then install the plugin: +Then install a plugin: ```shell codex plugin add docent@transluce-plugins +codex plugin add fxtr@transluce-plugins ``` diff --git a/plugins/fxtr/.codex-plugin/plugin.json b/plugins/fxtr/.codex-plugin/plugin.json new file mode 100644 index 0000000..747e6a9 --- /dev/null +++ b/plugins/fxtr/.codex-plugin/plugin.json @@ -0,0 +1,23 @@ +{ + "name": "fxtr", + "version": "0.1.0", + "description": "Write, run, and view fxtr experiments, and call models from them with behaviors.", + "author": { + "name": "Transluce", + "url": "https://transluce.org" + }, + "homepage": "https://docs.transluce.ai/fxtr", + "license": "Apache-2.0", + "keywords": ["fxtr", "behaviors", "experiments", "skills"], + "skills": "./skills/", + "interface": { + "displayName": "fxtr", + "shortDescription": "Write, run, and view fxtr experiments.", + "longDescription": "Skills for writing, running, and viewing fxtr experiments, and for calling models from them with behaviors.", + "developerName": "Transluce", + "category": "Productivity", + "capabilities": ["Read", "Write"], + "websiteURL": "https://docs.transluce.ai/fxtr", + "defaultPrompt": ["Use fxtr to set up a new experiment project.", "Use fxtr to sweep this prompt over several models and judge the replies."] + } +} diff --git a/plugins/fxtr/skills/behaviors/SKILL.md b/plugins/fxtr/skills/behaviors/SKILL.md new file mode 100644 index 0000000..9f81590 --- /dev/null +++ b/plugins/fxtr/skills/behaviors/SKILL.md @@ -0,0 +1,38 @@ +--- +name: behaviors +description: Call language models from fxtr experiments with the behaviors Python library, covering one-shot requests and judge steps, multi-turn chat rollouts, scripted or custom user and tool policies, retries and failures, streaming, checkpoints, and storing conversations. Use whenever a fxtr experiment calls a model or holds a conversation, or when extending behaviors' model adapters or policies. +--- + +# behaviors + +behaviors is a Python library for calling language models from fxtr experiments: provider +adapters, conversation sessions, context policies that play users and tools, retries and failure +recording, and conversation records stored as fxtr entities. Use these pieces when they fit the +experiment, rather than building another chat loop or transcript format. + +The pages below are its documentation. Read the ones a task needs before writing code. + +## Where to read + +| Task | Read | +| ------------------------------------------------------------------------------------------------------------------------------ | ---------------------------------------------------------------------------------- | +| Add behaviors to a project; hold a conversation, judge outputs, and summarize verdicts in fxtr steps; test without credentials | [Calling language models](references/guides/language-models.mdx) | +| Pick a provider, an adapter, and an exact model ID | [Choosing models](references/guides/choosing-models.mdx) | +| Script users, offer tools, or write a custom conversation flow | [Policies and tools](references/guides/policies-and-tools.mdx) | +| Rollout statuses and `max_turns`, failures, retries, streaming | [Rollout outcomes and retries](references/guides/rollout-outcomes.mdx) | +| Messages and content, stored records, transcripts, conversations in the viewer, checkpoints | [Conversation records and checkpoints](references/guides/conversation-records.mdx) | +| Run Docent readings, or convert records for Docent | [Docent readings](references/guides/docent-readings.mdx) | +| Choose a layer; imports, adapters, sampling settings, requests and responses | [API overview](references/reference/api.mdx) | +| Steps, workflows, arrays, launching jobs, and the step cache | the adjacent [fxtr skill](../fxtr/SKILL.md) | + +## Best practices + +* **Judge with `gpt-6-luna` by default.** For an LLM judge, use `gpt-6-luna` through the + `OpenAIResponsesAPI` adapter, unless the user asks for another model. +* **Keep the model the user names.** Verify its exact ID from the provider's listing, as Choosing + models shows, rather than substituting another. Without credentials to check, take candidate + IDs from the provider's published catalog, and tell the user that account access has not been + verified. +* **Shape model experiments as the guide does**: one step per conversation, with every setting + that affects the request in a configuration entity; fxtr replicas for repeated samples; and + separate steps for judging and for aggregating verdicts. diff --git a/plugins/fxtr/skills/behaviors/provenance.json b/plugins/fxtr/skills/behaviors/provenance.json new file mode 100644 index 0000000..0a46a01 --- /dev/null +++ b/plugins/fxtr/skills/behaviors/provenance.json @@ -0,0 +1,11 @@ +{ + "skill": "behaviors", + "documentation": "https://docs.transluce.ai/behaviors", + "packages": { + "fxtr": "0.0.1a0", + "behaviors": "0.0.1a0" + }, + "revision": "216737e51a92f5a0a83f7629af73d5a733fceb19", + "modified": false, + "content_sha256": "de0ca958d570464776b006b595d0f26b7be280a6c5baf88c46629d2982757eee" +} diff --git a/plugins/fxtr/skills/behaviors/references/INDEX.md b/plugins/fxtr/skills/behaviors/references/INDEX.md new file mode 100644 index 0000000..e5f4a36 --- /dev/null +++ b/plugins/fxtr/skills/behaviors/references/INDEX.md @@ -0,0 +1,19 @@ +# The behaviors documentation + +The pages of https://docs.transluce.ai/behaviors, copied with this skill: `provenance.json`, +beside its `SKILL.md`, says from which version. Each is listed with what it covers, +in the order of the site's navigation. + +## Get started + +- [What is the behaviors library?](index.mdx): Primitives for evaluating AI model behavior. + +## Other pages + +- [Choosing models](guides/choosing-models.mdx): Find the exact model IDs a provider offers your account, pick the adapter for its endpoint, and check a model before a sweep. +- [Conversation records and checkpoints](guides/conversation-records.mdx): What a rollout records, how conversations are stored as fxtr entities and flattened into transcripts, and how to checkpoint a session. +- [Docent readings](guides/docent-readings.mdx): Convert conversations to Docent's data models, and run a Docent reading against a transcript or agent run with a model you choose. +- [Calling language models](guides/language-models.mdx): Use the behaviors library to hold conversations and judge outputs in fxtr steps. +- [Policies and tools](guides/policies-and-tools.mdx): Decide what the model sees each turn: scripted users, tool policies, their combinations, and custom conversation flows. +- [Rollout outcomes and retries](guides/rollout-outcomes.mdx): How a rollout ends, how model call failures are retried, recorded, or raised, and how to watch a rollout as it runs. +- [API overview](reference/api.mdx): The layers of behaviors, where each is imported from, the provider adapters, and the request, response, and sampling types. diff --git a/plugins/fxtr/skills/behaviors/references/guides/choosing-models.mdx b/plugins/fxtr/skills/behaviors/references/guides/choosing-models.mdx new file mode 100644 index 0000000..c47cf47 --- /dev/null +++ b/plugins/fxtr/skills/behaviors/references/guides/choosing-models.mdx @@ -0,0 +1,110 @@ +--- +title: "Choosing models" +description: "Find the exact model IDs a provider offers your account, pick the adapter for its endpoint, and check a model before a sweep." +--- + +behaviors takes a provider's model ID as given: it has no catalog of models, and it passes the +string to the provider's endpoint without checking it against a list or translating it. Choose +the provider and endpoint first, then get the exact model ID from that provider's listing API. +Use the returned `id`, not a display name; the same model can have different IDs at different +providers. When you've been asked for a particular model, keep it, and verify its ID rather than +substituting another. + +The examples below use `curl` and `jq`, with API keys already set in the named environment +variables. They only list model metadata. Use the same credentials and endpoint as the +experiment, since a listing shows what that account can use. An authentication or permission +error does not mean that no models exist. + +## OpenAI + +[List models](https://developers.openai.com/api/reference/resources/models/methods/list) +with `GET https://api.openai.com/v1/models`: + +```bash +curl --fail-with-body --silent --show-error https://api.openai.com/v1/models \ + -H "Authorization: Bearer $OPENAI_API_KEY" \ + -o /tmp/openai-models.json && +jq -r '.data[].id' /tmp/openai-models.json | sort +``` + +Use a returned ID with `OpenAIResponsesAPI(model_id)` for models supporting the Responses API, or +`OpenAICompatibleAPI(model_id)` for Chat Completions. The listing includes models for other +tasks too; check a model's supported endpoints and features in the +[model catalog](https://developers.openai.com/api/docs/models) before choosing an adapter. + +## Anthropic + +[List models](https://platform.claude.com/docs/en/api/models/list) +with `GET https://api.anthropic.com/v1/models`: + +```bash +curl --fail-with-body --silent --show-error 'https://api.anthropic.com/v1/models?limit=1000' \ + -H "x-api-key: $ANTHROPIC_API_KEY" \ + -H 'anthropic-version: 2023-06-01' \ + -o /tmp/anthropic-models.json && +jq -r '.data[] | [.id, .display_name] | @tsv' /tmp/anthropic-models.json && +jq '{has_more, last_id}' /tmp/anthropic-models.json +``` + +This is one page, even with `limit=1000`. While `has_more` is true, request the next page with +`after_id` set to the returned `last_id`, and collect its models before advancing again. +Pass a returned `id` to `AnthropicMessagesAPI(model_id)`; `display_name` is only for presentation. + +## OpenRouter + +OpenRouter serves models from multiple vendors through an OpenAI-compatible endpoint. Its +[general catalog API](https://openrouter.ai/docs/api/api-reference/models/list-all-models-and-their-properties) +is `GET https://openrouter.ai/api/v1/models`; the +[account-filtered listing](https://openrouter.ai/docs/api/api-reference/models/list-models-filtered-by-user-provider-preferences-privacy-settings-and-guardrails) +also applies your provider preferences, privacy settings, and guardrails: + +```bash +curl --fail-with-body --silent --show-error https://openrouter.ai/api/v1/models/user \ + -H "Authorization: Bearer $OPENROUTER_API_KEY" \ + -o /tmp/openrouter-models.json && +jq -r '.data[] | [.id, .name] | @tsv' /tmp/openrouter-models.json +``` + +For the general catalog, change `/models/user` to `/models`. The +[model browser](https://openrouter.ai/models) is useful for interactive search. Inspect +`architecture`, `context_length`, and `supported_parameters` in the JSON for candidate models. +Use the full returned `id`, including its vendor prefix and any suffix, with: + +```python +import os + +from behaviors.implementations.models.openai_compatible import OpenAICompatibleAPI + +api = OpenAICompatibleAPI( + model_id, # Exact id selected from the OpenRouter listing. + base_url="https://openrouter.ai/api/v1", + api_key=os.environ["OPENROUTER_API_KEY"], +) +``` + +The OpenAI SDK does not read `OPENROUTER_API_KEY` by itself, so pass it explicitly as above. +Close the adapter with `await api.aclose()` when finished, as in the +[conversation step](language-models.mdx#a-conversation-as-a-step). + +## Search and validate a candidate + +Search a saved response by ID or display name, changing the file and search text as needed: + +```bash +jq -r --arg query 'claude' ' + .data[] + | select(([.id, (.name // .display_name // "")] | join(" ") | ascii_downcase) + | contains($query | ascii_downcase)) + | .id +' /tmp/openrouter-models.json +``` + +For paginated responses, search every collected page. A catalog entry does not establish that +the model supports every feature the experiment needs. Check its support for the endpoint, +streaming, tools, and reasoning, then try the chosen adapter and sampling settings with a small +request before launching a sweep. Without credentials, the providers' published catalogs and +documentation give candidate IDs, but not whether your account can use them. + +Keep the chosen provider, model ID, and request settings in the experiment's configuration +inputs, so that each run records what was selected (see +[A conversation as a step](language-models.mdx#a-conversation-as-a-step)). diff --git a/plugins/fxtr/skills/behaviors/references/guides/conversation-records.mdx b/plugins/fxtr/skills/behaviors/references/guides/conversation-records.mdx new file mode 100644 index 0000000..cb317c2 --- /dev/null +++ b/plugins/fxtr/skills/behaviors/references/guides/conversation-records.mdx @@ -0,0 +1,81 @@ +--- +title: "Conversation records and checkpoints" +description: "What a rollout records, how conversations are stored as fxtr entities and flattened into transcripts, and how to checkpoint a session." +--- + +A rollout returns a `ChatRollout`: its context turns, its model turns, how it ended, and the +initialization its model session started from. This page covers what those records hold, how they +are stored and shown, and how they differ from a session's checkpoint. + +## Messages and content + +Messages are `SystemMessage`, `UserMessage`, `AssistantMessage`, and `ToolMessage`, each holding a +tuple of typed content blocks. An assistant's content may include text, tool calls, reasoning, +redacted reasoning, or summarized reasoning, so don't assume every block has `.text`. For the +answer's text alone, keep the `ContentText` blocks of the turn's `AssistantMessage`s, as the +[judge step's](language-models.mdx#a-judge-step) `answer_text` does, and keep the +structured rollout for later inspection. + +A model turn made by an API-backed session records the provider's response in +`ModelTurn.metadata["response"]`, as JSON, including its stop reason and usage. Metadata is defined +by whatever produced the turn, so validate it before scoring with it. + +## Storing records + +`ChatRollout`, `Transcript`, `TranscriptGroup`, `AgentRun`, and the concrete session +initializations, such as `SetSystemMessage`, are fxtr entities, with type IDs of the form +`org.transluce.behaviors..v1`. A step that returns a `ChatRollout` has it stored like any +other entity (see [Entities and custom types](../../../fxtr/references/concepts/entities-and-custom-types.mdx)), and +`await storage.store(entity)` and `await storage.load(ref)` work on them too. To load a bare ID, +pass the concrete type, as in `as_type=ChatRollout`, not an abstract base such as +`SessionInitialization`. A tree of transcripts and groups is stored as one entity, so there is no +need to store each child separately. + +fxtr checks stored records strictly: give fields their declared types (`temperature=0.0`, not +`0`), and keep metadata to JSON values with string keys. + +## Transcripts + +`ChatRolloutTranscriptProjection`, from `behaviors.rollouts.chat`, flattens a rollout into a plain +`Transcript`: `await ChatRolloutTranscriptProjection().to_transcript(rollout)` gives one message +list, starting with the system message of a recorded `SetSystemMessage` initialization and +followed by the turns in order, including a final context turn the model never answered. +`to_agent_run` wraps that transcript in an `AgentRun`. The projection drops each turn's metadata, +so judge from the native rollout when you need usage, stop reasons, or other metadata. + +## Showing conversations in the viewer + +The `fxtr-view-behaviors` package draws rollouts, transcripts, transcript groups, agent runs, and +system-message initializations as conversations. To use it, spread its `behaviorsRenderers` into +your project's bundle: + +```ts views/src/bundle.ts +import type { ViewBundle } from "fxtr-view"; +import { behaviorsRenderers } from "fxtr-view-behaviors"; + +const bundle: ViewBundle = { + renderers: [...behaviorsRenderers], +}; + +export default bundle; +``` + +A new project's renderer package doesn't depend on it yet: add +`"fxtr-view-behaviors": "link:./vendor/fxtr/packages/fxtr-view-behaviors"` to the dependencies in +`views/package.json`, like the other fxtr packages, and run `pnpm install` in `views/` (see +[The renderer package](../../../fxtr/references/reference/renderers.mdx#the-renderer-package)). The package also exports +`ChatRolloutView`, for a rollout inside a view of your own, and `ComparedRollouts`. + +## Checkpointing sessions + +A serializable session can save its state and be restored from it. `await session.get_state()` +returns a JSON object, and `await factory.restore_session(state)` starts a session in that state. +The sessions of `chat_session_model_from_api` and of `ScriptedUserSimulator` support this, and the +`serializable_context_policy_from_*` combinators build checkpointable policies when every component +they combine is checkpointable. Generator-based policies can't be checkpointed, because a +suspended generator can't be saved. + +A checkpoint is not a stored record: it is a session's working state, apart from the transcript +entities and from fxtr's step cache. fxtr doesn't resume a partly finished rollout from a +checkpoint: a step that failed partway runs again from the start when its job resumes. Close the +sessions you start yourself with `aclose()`. diff --git a/plugins/fxtr/skills/behaviors/references/guides/docent-readings.mdx b/plugins/fxtr/skills/behaviors/references/guides/docent-readings.mdx new file mode 100644 index 0000000..e372d8f --- /dev/null +++ b/plugins/fxtr/skills/behaviors/references/guides/docent-readings.mdx @@ -0,0 +1,47 @@ +--- +title: "Docent readings" +description: "Convert conversations to Docent's data models, and run a Docent reading against a transcript or agent run with a model you choose." +--- + +Docent is Transluce's platform for analyzing AI agent behavior, whose readings ask a language model +questions about transcripts. behaviors can convert its conversation records to Docent's data +models, and run a Docent reading locally: +render a reading's prompt against a transcript or agent run, call a model through a `ModelAPI` +you provide, and parse and validate the answer and its citations. Install the `docent` extra, +`behaviors[docent]`, and import from `behaviors.implementations.docent`. + +## Converting records + +`to_docent_transcript(transcript)` and `to_docent_agent_run(agent_run)` turn behaviors' +`Transcript` and `AgentRun` (see +[Transcripts](conversation-records.mdx#transcripts)) into the Docent SDK's objects, +and `to_docent_message(message)` converts one message. + +## Running a reading + +`DocentReadingStepConfig` describes one reading: + +* `prompt_template_segments`: the segments of a Docent reading's prompt template, a tuple of + strings and placeholder dictionaries, such as + `{"param_name": "conversation", "param_type": "transcript", "is_list": False}`; +* `output_schema`: the JSON Schema the answer must conform to; +* `input_param_name`: the placeholder the transcript or agent run fills, which must be one of + the segments' placeholders (`param_type` `"transcript"` for a `Transcript`, `"agent_run"` for an + `AgentRun`); +* `max_output_attempts`: how many times to sample the model when its answer doesn't parse or + validate, 3 by default; +* `default_parsed_response_on_failure`: what to return when those attempts run out; without it, + the last validation error is raised. + +Then construct `DocentReadingStep(config=config, model_api=api, sampling=sampling)` and call +`await reader.read(transcript_or_agent_run)` with an input of the type the template declares. +`retry=RetryConfig(...)` retries transport failures, as for any model call (see +[Failures and retries](rollout-outcomes.mdx#failures-and-retries)), separately from +the attempts at a valid answer. + +`read` returns a `DocentReadingStepResult` with the `parsed_response`, the `reader_transcript` of +the model call, and the `retry_error_transcripts` of attempts whose answers failed to validate. The +fallback from `default_parsed_response_on_failure` has no reader transcript. A transport failure +that outlasts its retries is returned as a `ModelCallFailure`. You own the model API you pass in: +close it when the reading is done, as the +[judge step](language-models.mdx#a-judge-step) closes its adapter. diff --git a/plugins/fxtr/skills/behaviors/references/guides/language-models.mdx b/plugins/fxtr/skills/behaviors/references/guides/language-models.mdx new file mode 100644 index 0000000..2e9af1e --- /dev/null +++ b/plugins/fxtr/skills/behaviors/references/guides/language-models.mdx @@ -0,0 +1,337 @@ +--- +title: "Calling language models" +description: "Use the behaviors library to hold conversations and judge outputs in fxtr steps." +--- + +**behaviors** is a Python library for calling language models from fxtr experiments. It provides: + +* model adapters for Anthropic, OpenAI, and OpenAI-compatible services; +* conversation sessions and **context policies**, which decide what the model sees next: + scripted user messages, tool results, or custom logic; +* retries for rate limits and server errors; +* conversation records that are fxtr entities, so they're stored, linked, and shown in the + viewer. + +## Set up + +Add behaviors, with its provider SDKs, to your project: + +```bash +uv add "behaviors[models]" +``` + +Then commit the updated `pyproject.toml` and `uv.lock`. + +Set your provider's API key, `OPENAI_API_KEY` or `ANTHROPIC_API_KEY`, in the project's `.env` +(see [Project configuration](../../../fxtr/references/reference/project-configuration.mdx#env)) or in the environment. + +| Provider | Adapter | +| ---------------------------------- | ------------------------------------------------------------------------------------ | +| OpenAI Responses | `from behaviors.implementations.models.openai_responses import OpenAIResponsesAPI` | +| Anthropic Messages | `from behaviors.implementations.models.anthropic import AnthropicMessagesAPI` | +| OpenAI Chat Completions-compatible | `from behaviors.implementations.models.openai_compatible import OpenAICompatibleAPI` | + +Construct an adapter with the provider's model ID, such as `OpenAIResponsesAPI("gpt-6-luna")`. +[Choosing models](choosing-models.mdx) shows how to find the IDs your account can +use, and the [API overview](../reference/api.mdx#provider-adapters) covers the adapters' +other options, such as `base_url` for a compatible service. + +## A conversation as a step + +This step holds a two-turn conversation: it asks a question, then asks the model to explain its +answer. The model and sampling settings arrive as a configuration entity, so every conversation +points at the exact settings that produced it. Keep every setting that affects the request in the +configuration, so that changing one changes the step's inputs. + +```python +from dataclasses import dataclass +from typing import Annotated, cast + +import anyio + +from behaviors.implementations.context_policies import ( + ScriptedUserSimulator, + context_policy_from_user_simulator, +) +from behaviors.implementations.models import chat_session_model_from_api +from behaviors.implementations.models.openai_responses import OpenAIResponsesAPI +from behaviors.rollouts.chat import run_chat_rollout +from behaviors.types import ( + ChatRollout, ContentText, RetryConfig, SamplingParams, SetSystemMessage, + SystemMessage, UserMessage, +) +from fxtr.core.entities import BoundID +from fxtr.entity_defns.dataclass_entity import DataclassEntity +from fxtr.experiment.handles import ArrayHandle +from fxtr.experiment.steps import StepContext, step +from fxtr.experiment.workflows import WorkflowContext, workflow + + +@dataclass(frozen=True) +class ConversationConfig(DataclassEntity, fxtr_type="org.example.my_experiment.ConversationConfig.v1"): + model_name: str + sampling: SamplingParams + system_prompt: str + + +def make_api(model_name: str) -> OpenAIResponsesAPI: + return OpenAIResponsesAPI(model_name) + + +@step(name="my_experiment.conversation") +async def converse(context: StepContext, config: ConversationConfig, prompt: str) -> ChatRollout: + """Answer a prompt, then explain the answer in a follow-up turn.""" + api = make_api(config.model_name) + try: + model = chat_session_model_from_api(api, config.sampling, retry=RetryConfig()) + user = ScriptedUserSimulator(messages=( + UserMessage(content=(ContentText(text=prompt),)), + UserMessage(content=(ContentText(text="Explain your answer."),)), + )) + policy = context_policy_from_user_simulator( + user, + initialization=SetSystemMessage( + message=SystemMessage(content=(ContentText(text=config.system_prompt),)), + ), + ) + # One more turn than there are scripted prompts, so the rollout can end as completed. + return await run_chat_rollout(model, policy, max_turns=3) + finally: + with anyio.CancelScope(shield=True): + await api.aclose() + + +@workflow(name="my_experiment.run") +async def experiment( + context: WorkflowContext, + configs: Annotated[ArrayHandle[BoundID[ConversationConfig]], "[model: str]"], + prompts: Annotated[ArrayHandle[str], "[prompt: str]"], +) -> Annotated[ArrayHandle[BoundID[ChatRollout]], "[model: str, prompt: str]"]: + """Hold one conversation per model and prompt.""" + rollouts = context.run_step( + "conversations", + converse, + {"config": configs, "prompt": prompts}, + map_over=["model", "prompt"], + ) + return cast(ArrayHandle[BoundID[ChatRollout]], rollouts) +``` + +In this step: + +* `chat_session_model_from_api` wraps the adapter as a conversation model. Pass + `retry=RetryConfig()` to retry rate limits, overloads, and server errors; without it, calls + aren't retried. +* `ScriptedUserSimulator` sends one scripted message per turn and then ends the conversation. + `context_policy_from_user_simulator` turns it into a context policy, here starting with a system + message. +* `run_chat_rollout` alternates between the policy and the model, and returns a `ChatRollout`: + the whole conversation, with its status. The step's output is an entity, so fxtr stores it. +* Building the adapter in one helper, `make_api`, lets you substitute a fake model for + credential-free checks. +* The adapter is closed in `finally`, because `run_chat_rollout` closes the conversation + sessions but not the provider client. + +Put the step and workflow in a module of your project, here `my_project.conversations`, and list +it in `[tool.fxtr] modules` (see +[Project configuration](../../../fxtr/references/reference/project-configuration.mdx)). Launch it with a stored +configuration per model: + +```python +import anyio + +from behaviors.types import SamplingParams +from fxtr.core.entities import BoundID +from fxtr.entity_defns.array import Array +from fxtr.project.running import open_local_client, run_job +from my_project.conversations import ConversationConfig, experiment + + +async def main() -> None: + async with open_local_client(__file__) as client: + luna = await client.store(ConversationConfig( + model_name="gpt-6-luna", + sampling=SamplingParams(max_tokens=2000), + system_prompt="Answer concisely.", + )) + configs = Array.from_items([(("luna",), luna)], ("[model: str]", BoundID[ConversationConfig])) + prompts = Array.from_items([(("greeting",), "Hello!")], ("[prompt: str]", str)) + await run_job(client, experiment, {"configs": configs, "prompts": prompts}, root="conv-v1") + + +if __name__ == "__main__": + anyio.run(main) +``` + +For repeated samples of each conversation, add `replicas={"sample": range(n)}` to the +`run_step` call (see [Replicas](../../../fxtr/references/concepts/arrays-and-parallel-computations.mdx#replicas)). + +## A judge step + +Judging needs no conversation state, so a judge calls the model once with `generate_with_retries`. +This step grades a conversation's final answer and returns a `Verdict` entity, which the viewer +can link to. Put it in the same module as the conversation step, whose `make_api` it uses: + +```python +import re +from dataclasses import dataclass +from typing import Annotated + +import anyio + +from behaviors.implementations.models import ( + DEFAULT_FAILURE_CATEGORIES_TO_RETRY, + generate_with_retries, +) +from behaviors.types import ( + AssistantMessage, ChatMessage, ChatRollout, ContentText, ModelCallFailure, + ModelRequest, RetryConfig, RolloutCompleted, SamplingParams, UserMessage, +) +from fxtr.core.entities import BoundID +from fxtr.entity_defns.array import Array +from fxtr.entity_defns.dataclass_entity import DataclassEntity +from fxtr.experiment.steps import StepContext, step + + +@dataclass(frozen=True) +class JudgeConfig(DataclassEntity, fxtr_type="org.example.my_experiment.JudgeConfig.v1"): + model_name: str + sampling: SamplingParams + rubric: str + + +@dataclass(frozen=True) +class Verdict(DataclassEntity, fxtr_type="org.example.my_experiment.Verdict.v1"): + score: int | None # None when there was nothing to grade or the reply had no score + judge_reply: str + + +def answer_text(messages: tuple[ChatMessage, ...]) -> str: + """The assistant's text, without reasoning or tool-use blocks.""" + return "".join( + block.text + for message in messages + if isinstance(message, AssistantMessage) + for block in message.content + if isinstance(block, ContentText) + ) + + +@step(name="my_experiment.judge") +async def judge(context: StepContext, judge_config: JudgeConfig, rollout: ChatRollout) -> Verdict: + """Grade the conversation's final answer against the rubric from 0 to 10.""" + if not isinstance(rollout.status, RolloutCompleted) or not rollout.model_turns: + return Verdict(score=None, judge_reply=f"not graded: {rollout.status.kind}") + prompt = ( + f"{judge_config.rubric}\n\n" + f"Answer to grade:\n{answer_text(rollout.model_turns[-1].messages)}\n\n" + "End your reply with a line of the form 'Score: N', where N is from 0 to 10." + ) + request = ModelRequest( + context=(UserMessage(content=(ContentText(text=prompt),)),), + tools=(), + tool_choice=None, + parallel_tool_calls=None, + sampling=judge_config.sampling, + ) + api = make_api(judge_config.model_name) + try: + response = await generate_with_retries( + api, + request, + retry=RetryConfig(), + failure_categories_to_retry=DEFAULT_FAILURE_CATEGORIES_TO_RETRY, + ) + finally: + with anyio.CancelScope(shield=True): + await api.aclose() + if isinstance(response, ModelCallFailure): + # Raise rather than return: a returned failure would be cached as the result. + raise RuntimeError(f"judge call failed ({response.category}): {response.description}") + reply = answer_text(response.output) + match = re.search(r"Score:\s*(\d+)\s*$", reply) + score = int(match.group(1)) if match else None + return Verdict(score=score if score is not None and 0 <= score <= 10 else None, judge_reply=reply) +``` + +Schedule it in the workflow, mapped over the conversations, with the judge configuration as a +single stored input. The `judge_config: JudgeConfig` and `rollout: ChatRollout` annotations load +the entities before the body runs: + +```python +verdicts = context.run_step( + "judge_final_answers", + judge, + {"judge_config": judge_config, "rollout": rollouts}, + map_over=["model", "prompt"], +) +``` + +The judge follows the [return or raise](../../../fxtr/references/concepts/workflows-and-steps.mdx#failures-return-or-raise) +rule: it returns a `Verdict` for outcomes every attempt would reproduce, such as a conversation +that never completed or a reply without a score. It raises when the model call still fails after +its retries, so resuming the job calls the judge again. + +## Summarize the verdicts + +A reduction step takes one model's verdict references and loads them: + +```python +@step(name="my_experiment.mean_score") +async def mean_score( + context: StepContext, + verdicts: Annotated[Array[BoundID[Verdict]], "[prompt: str]"], +) -> float | None: + """Average the judge's scores across prompts.""" + scores: list[int] = [] + for ref in verdicts.values(): + verdict = await context.load(ref) + if verdict.score is not None: + scores.append(verdict.score) + return sum(scores) / len(scores) if scores else None + + +# In the workflow: +mean_scores = context.run_step("mean_scores", mean_score, {"verdicts": verdicts}, map_over=["model"]) +``` + +## Rollouts, settings, and testing + + + + A rollout ends as `RolloutCompleted`, `RolloutMaxTurns`, or `RolloutFailed`. `max_turns` + counts model turns: a script with N prompts needs `max_turns=N + 1` to end as completed. A + model stop reason such as `StopMaxTokens` is separate from the rollout's status, and may mean + the content was truncated even in a completed rollout. See + [How a rollout ends](rollout-outcomes.mdx#how-a-rollout-ends). + + + + By default, a rollout that exceeds the model's context length is returned as `RolloutFailed`, + keeping the conversation so far; other failures raise. A returned failed rollout is a + successful step result, so it's cached. Decide deliberately how scoring treats such outcomes. + See [Recorded failures](rollout-outcomes.mdx#recorded-failures). + + + + `SamplingParams` covers `max_tokens`, `temperature`, `top_p`, `stop_sequences`, `reasoning`, + and provider-specific `extra`. Which settings take effect depends on the adapter and model: + for example, `OpenAIResponsesAPI` doesn't apply `stop_sequences`. Leave a setting unset + unless your adapter and model support it. See + [Sampling settings](../reference/api.mdx#sampling-settings). + + + + Any object with an async `generate(request)` method returning a `ModelResponse`, and an + `aclose()` method, can stand in for a provider adapter. Return it from `make_api` in tests, and + run against a [scratch schema](../../../fxtr/references/guides/running-and-viewing-jobs.mdx#before-you-launch). + + + + `context_policy_from_user_and_tools` combines a user simulator with a tool policy you + implement, which defines the tools and executes their calls. For fully custom flows, + `context_policy_from_generator` turns an async generator into a context policy: it yields each + turn's messages and receives the model's reply. See + [Policies and tools](policies-and-tools.mdx). + + diff --git a/plugins/fxtr/skills/behaviors/references/guides/policies-and-tools.mdx b/plugins/fxtr/skills/behaviors/references/guides/policies-and-tools.mdx new file mode 100644 index 0000000..923316d --- /dev/null +++ b/plugins/fxtr/skills/behaviors/references/guides/policies-and-tools.mdx @@ -0,0 +1,84 @@ +--- +title: "Policies and tools" +description: "Decide what the model sees each turn: scripted users, tool policies, their combinations, and custom conversation flows." +--- + +A **context policy** is the model's environment in a rollout. It supplies the session's +initialization, such as a system message, then each turn the model sees next (a `ContextTurn`), +until it ends the rollout with `EndRollout`. `run_chat_rollout(model, policy, ...)` alternates the +policy's turns with the model's. behaviors builds policies from smaller parts: a **user +simulator** writes the user's messages, and a **tool policy** answers the model's tool calls. +These come from `behaviors.implementations.context_policies`. + +## Scripted users + +`ScriptedUserSimulator(messages=(...))` sends one scripted message each time it is consulted, +then returns `EndRollout()`. `context_policy_from_user_simulator(user, initialization=...)` makes +it a policy, as in the [conversation step](language-models.mdx#a-conversation-as-a-step). +That policy never offers the model tools, and raises if the model calls one anyway. + +## Tools + +Tool definitions describe tools to the model; behaviors never runs a function because a tool +call names it. The tool policy executes calls and validates their inputs. A tool policy implements +these async methods, whose structural protocols are `ToolPolicy` and `ToolPolicySession` in +`behaviors.interfaces.context_policies`: + +* `ToolPolicy.start_session()` creates an independent tool session for one rollout. +* `available_tools()` returns the tool definitions: `ToolParam`s, each with a name, a + description, a JSON schema of its input, `strict`, and `input_examples` (pass `None` explicitly + for an optional field you don't use). +* `generate_tool_turn(tool_uses)` returns a `ToolTurn` with one `ToolMessage` for each + `ContentToolUse`, in the same order and with matching `tool_call_id`s. +* `aclose()` releases the session's resources. + +`NoTools` is a tool policy that offers none. + +Two combinators make a tool policy part of a context policy: + +* `context_policy_from_user_and_tools(user, tools, initialization=...)` sends model turns that + contain tool calls to the tool policy, and every other turn to the user simulator, which sees + the model's messages since it was last consulted. By default + (`show_tool_calls_to_user=False`), the simulated user doesn't see tool calls or their results. +* `context_policy_from_tool_policy(tools, first_turn_messages=..., initialization=...)` starts + with fixed messages, answers each model turn's tool calls, and ends the rollout when the model + takes a turn without tool calls. Cap a runaway loop with `run_chat_rollout`'s `max_turns`. + +A tool policy can't end a rollout; only a user simulator or a full context policy can. An +environment that should end when, say, a submit tool completes a task needs a context policy of +its own (below). + +## Custom conversation flows + +For a flow the combinators don't express (a user who answers depending on what the model said, +tools offered only at some turns), write the policy as an async generator: each `yield` gives the +next context turn, and the model's reply comes back as the value of the `yield`: + +```python +from behaviors.implementations.context_policies import ( + ContextPolicyGenerator, context_policy_from_generator, +) +from behaviors.types import ContentText, ContextTurn, EndRollout, UserMessage + +async def turns() -> ContextPolicyGenerator: + reply = yield ContextTurn( + messages=(UserMessage(content=(ContentText(text="Choose a color."),)),), + available_tools=None, + ) + # Inspect reply.messages here to choose a data-dependent next turn if needed. + yield EndRollout(metadata={"received_reply": reply is not None}) + +policy = context_policy_from_generator(turns, initialization=None) +``` + +Pass the generator function (`turns`), not a running generator: each rollout starts its own. +Yield `EndRollout` explicitly to finish, since a generator that simply returns raises an error. +For a reusable policy with settings of its own, subclass `GeneratorContextPolicy`, overriding +`generate_turns`, and `generate_initialization` if its sessions need one. + +A generator-based policy can't be checkpointed, because a suspended generator can't be saved. A +policy that must be checkpointed implements the `ContextPolicy` and `ContextPolicySession` +protocols directly (see +[Checkpointing sessions](conversation-records.mdx#checkpointing-sessions)). The +combinators above are one reasonable way to sequence users and tools; a flow with different rules +can be written the same way, with their source as a reference. diff --git a/plugins/fxtr/skills/behaviors/references/guides/rollout-outcomes.mdx b/plugins/fxtr/skills/behaviors/references/guides/rollout-outcomes.mdx new file mode 100644 index 0000000..2800ddb --- /dev/null +++ b/plugins/fxtr/skills/behaviors/references/guides/rollout-outcomes.mdx @@ -0,0 +1,77 @@ +--- +title: "Rollout outcomes and retries" +description: "How a rollout ends, how model call failures are retried, recorded, or raised, and how to watch a rollout as it runs." +--- + +A rollout can end in more ways than a completed conversation, and a model call can fail. This page +covers what a returned `ChatRollout` says about how it ended, how behaviors retries and records +failures, and what that means for a fxtr step that returns the rollout. + +## How a rollout ends + +A rollout's `status` is one of: + +* `RolloutCompleted`: the context policy ended the rollout with `EndRollout`, which is recorded + on the rollout's `end`. +* `RolloutMaxTurns`: the rollout reached `run_chat_rollout`'s `max_turns`. +* `RolloutFailed`: a model call failed in a category the rollout records (see + [Recorded failures](#recorded-failures)). + +Check the status before treating a rollout as a successful observation. `max_turns` counts +completed **model** turns, and the limit is checked before the policy is asked for its next turn. +So a script with exactly N prompts and `max_turns=N` ends as `RolloutMaxTurns` without the policy +being asked for its final `EndRollout`: allow one more turn, `max_turns=N + 1`, when how the policy +ends matters. With `max_turns=None`, the default, the runner imposes no limit. + +A model turn's stop reason, such as `StopMaxTokens`, is separate from the rollout's status: a +completed rollout can still contain a reply cut off at the token limit. + +## Failures and retries + +A model call's expected failures are returned as data: `ModelAPI.generate` returns a +`ModelCallFailure` with a `category` (`context_length`, `content_filter`, `rate_limited`, +`overloaded`, `authentication`, `invalid_request`, or `server_error`), a `description`, and the +provider's HTTP status. An error in the implementation itself raises. The provider adapters turn +off their SDKs' own retries, so that these settings decide every attempt: + +* A session of `chat_session_model_from_api` has **no retry schedule by default** (`retry=None`). + Pass `retry=RetryConfig()` to retry failures in `failure_categories_to_retry`, by default + `DEFAULT_FAILURE_CATEGORIES_TO_RETRY`: `rate_limited`, `overloaded`, and `server_error`. +* For one request, `generate_with_retries(api, request, retry=..., failure_categories_to_retry=...)` + takes the same settings, both required, as in the + [judge step](language-models.mdx#a-judge-step). Once attempts run out, it returns + the last failure. + +`RetryConfig` sets the schedule: `max_attempts` (3 by default, counting the first), +`initial_backoff_seconds` (1), `backoff_base` (3), `max_backoff_seconds` (120), and +`jitter_seconds` (up to 5 seconds added to every wait at random). + +## Recorded failures + +When a session's model call still fails after its retries, `run_chat_rollout` either records the +failure or raises it. By default, `failure_categories_to_record=("context_length",)`: a rollout +whose context outgrows the model's window is returned as `RolloutFailed`, with the conversation so +far, including the context turn the model never answered. A failure in any other category raises +`ModelCallFailedError`, unless you add its category to `failure_categories_to_record`. + +A recorded failure is a returned value, so a fxtr step that returns the rollout succeeds, and fxtr +caches the failed rollout as the step's result (see +[Failures: return or raise?](../../../fxtr/references/concepts/workflows-and-steps.mdx#failures-return-or-raise)). Record +a category only for failures every attempt would reproduce, and decide deliberately how scoring +treats them, as missing or as failed, rather than scoring them as successes. + +## Watching a rollout as it runs + +Pass an async `event_callback` to `run_chat_rollout` to receive the rollout's events as they +happen: + +* `ContextTurnEvent` and `ModelTurnEvent`, each carrying a completed turn, the same object that + lands on the returned rollout; +* `ModelStreamEvent`, carrying a delta of the model's output while it generates a turn, when the + model's session streams (sessions of a streaming adapter, such as the provider adapters, do). + +The callback is awaited in order, on the rollout's own path: a slow callback slows the rollout, +and one that raises aborts it. Deltas can describe output that a retry then discards: +`StreamRestart` tells a consumer to reset what it has assembled of the turn so far. The completed +turns and the returned rollout are the durable content, and how the rollout ended is its returned +status, not an event. diff --git a/plugins/fxtr/skills/behaviors/references/index.mdx b/plugins/fxtr/skills/behaviors/references/index.mdx new file mode 100644 index 0000000..0108759 --- /dev/null +++ b/plugins/fxtr/skills/behaviors/references/index.mdx @@ -0,0 +1,40 @@ +--- +title: "What is the behaviors library?" +description: "Primitives for evaluating AI model behavior." +--- + + + The behaviors library documentation is still under construction. + + +**behaviors** is a Python library for building AI evaluations. It provides model adapters, +multi-turn conversations, user and tool policies, and structured conversation records. +Use it to sample a model, simulate an interaction, or ask a model to judge another model's +output. + +You control the messages, tools, and stopping conditions of an interaction. behaviors handles +the conversation loop and provides configurable retries and failure recording. It supplies no +general-purpose judge and no model-driven user simulator: you write those for your experiment +from its interfaces, as the [judge step](guides/language-models.mdx#a-judge-step) does. + +behaviors works with [fxtr](../../fxtr/references/index.mdx): conversation records can be stored as fxtr entities, +returned from steps, and inspected in the experiment viewer. fxtr manages the experiment's +execution, caching, and provenance. For the shortest example of a step that calls a model, see +[Steps](../../fxtr/references/concepts/workflows-and-steps.mdx#steps) in the fxtr concepts. + + + + Set up behaviors, run conversations, and judge their outputs in fxtr steps. + + + + Set up the project that runs and stores your experiments. + + + +The other guides go further: [choosing models](guides/choosing-models.mdx), +[policies and tools](guides/policies-and-tools.mdx), +[rollout outcomes and retries](guides/rollout-outcomes.mdx), +[conversation records and checkpoints](guides/conversation-records.mdx), and +[Docent readings](guides/docent-readings.mdx). The +[API overview](reference/api.mdx) maps the library's layers and types. diff --git a/plugins/fxtr/skills/behaviors/references/reference/api.mdx b/plugins/fxtr/skills/behaviors/references/reference/api.mdx new file mode 100644 index 0000000..0afd75c --- /dev/null +++ b/plugins/fxtr/skills/behaviors/references/reference/api.mdx @@ -0,0 +1,110 @@ +--- +title: "API overview" +description: "The layers of behaviors, where each is imported from, the provider adapters, and the request, response, and sampling types." +--- + +behaviors is built in layers, from one model request up to a whole conversation with an +environment. Use the lowest layer that fits: a judge needs one request, a conversation needs a +session, and a conversation against simulated users or tools needs a context policy. The +[guide to calling language models](../guides/language-models.mdx) shows the common cases as +fxtr steps. + +## Choose a layer + +| Need | API | +| -------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------- | +| One model request, no conversation state | `ModelAPI.generate(ModelRequest)`, from `behaviors.interfaces.model_apis`; returns a `ModelResponse` or a `ModelCallFailure`. | +| Incremental output from a request | `StreamingModelAPI.generate_with_events(request, callback)`. | +| One request with a retry schedule | `generate_with_retries(api, request, retry=..., failure_categories_to_retry=...)`, from `behaviors.implementations.models`. | +| A conversation with accumulated messages | `chat_session_model_from_api(api, sampling, retry=...)`, from `behaviors.implementations.models`, which creates an independent session per conversation. | +| A model against an environment or a simulated user | `run_chat_rollout(model, context_policy, ...)`, from `behaviors.rollouts.chat`; returns a `ChatRollout`. | +| Fixed user prompts | `ScriptedUserSimulator(messages=(...))`, made a policy with `context_policy_from_user_simulator(..., initialization=...)`. | +| User prompts interleaved with tool calls | `context_policy_from_user_and_tools(user, tools, initialization=...)`. | +| A first turn followed only by tool turns | `context_policy_from_tool_policy(tools, first_turn_messages=..., initialization=...)`. | +| A custom or adaptive conversation | `context_policy_from_generator`, or a subclass of `GeneratorContextPolicy`; for checkpoints, implement the `ContextPolicy` and session protocols directly. | +| Plain transcripts for viewers or readers | `ChatRolloutTranscriptProjection`, from `behaviors.rollouts.chat`, with `Transcript`, `TranscriptGroup`, and `AgentRun`. | +| Docent readings | `DocentReadingStep` and the conversions in `behaviors.implementations.docent`, with `behaviors[docent]`. | + +The guides cover the layers above the request: +[policies and tools](../guides/policies-and-tools.mdx), +[rollout outcomes and retries](../guides/rollout-outcomes.mdx), +[conversation records and checkpoints](../guides/conversation-records.mdx), and +[Docent readings](../guides/docent-readings.mdx). behaviors supplies no general judge and no +model-driven user simulator: build those for your experiment from these interfaces, as the +[judge step](../guides/language-models.mdx#a-judge-step) does. + +## Where things are imported from + +* `behaviors.types` exports the data types: messages and content blocks, requests, responses, + sampling and reasoning parameters, failures, retry settings, stream events, rollouts and their + turns and statuses, and transcripts. +* `behaviors.implementations.models` exports `chat_session_model_from_api`, + `generate_with_retries`, `DEFAULT_FAILURE_CATEGORIES_TO_RETRY`, and `SamplingOverrides`. +* `behaviors.implementations.context_policies` exports the user simulators, tool policies, and + policy combinators. +* `behaviors.interfaces` holds the protocols: `model_apis`, `chat_session_models`, + `context_policies`, and `projections`. +* The provider adapters are each in a module of their own (below), imported explicitly, because + they need the provider SDKs. + +## Sessions and factories + +`ModelAPI` keeps no state between calls: every request carries the whole context. A +`ChatSessionModel` is a factory whose sessions (`ChatSession`) keep a conversation's history, and +a `ContextPolicy` is a factory whose sessions supply the model's initialization and then each next +`ContextTurn`, or `EndRollout` to finish. Keep mutable state for one conversation in its sessions +(or a generator's locals), never on a factory that every conversation shares. + +`run_chat_rollout` starts a session of each factory and closes both when it returns, including on +cancellation. It does not own the provider client, which the caller closes with +`await api.aclose()`. Use a fresh session for each independent conversation. The sessions of +`chat_session_model_from_api` accept `None` or `SetSystemMessage` as their initialization; any +other kind needs a session model that understands it. Close sessions you start yourself with +`aclose()`. + +## Provider adapters + +Install the provider SDKs with the `models` extra, `behaviors[models]`. The core types, context +policies, and rollouts with a fake model need no SDK. + +| Provider interface | Adapter | +| ------------------------------------------------ | ------------------------------------------------------------------------------------ | +| Anthropic Messages | `from behaviors.implementations.models.anthropic import AnthropicMessagesAPI` | +| OpenAI Responses | `from behaviors.implementations.models.openai_responses import OpenAIResponsesAPI` | +| OpenAI Chat Completions, and compatible services | `from behaviors.implementations.models.openai_compatible import OpenAICompatibleAPI` | + +Construct an adapter with the provider's model ID, as in `OpenAIResponsesAPI("gpt-6-luna")`; +[Choosing models](../guides/choosing-models.mdx) shows how to find IDs. Adapters read +credentials from the environment (`OPENAI_API_KEY`, `ANTHROPIC_API_KEY`). Each also takes +`api_key` and `base_url`, for another account or a compatible service, or an SDK client of your +own with `client=...`, but not both. An adapter closes a client it created, and leaves an +injected client to its caller. Never store credentials in experiment inputs or entities. + +The adapters turn off the SDKs' own retries, so that behaviors decides when a call is tried again +(see [Rollout outcomes and retries](../guides/rollout-outcomes.mdx#failures-and-retries)). + +## Requests and responses + +A `ModelRequest` has `context` (the messages), `tools`, `tool_choice`, `parallel_tool_calls`, and +`sampling`, all required: pass `()` or `None` where they don't apply. `generate` returns a +`ModelResponse` with `id`, `output` (a tuple of messages), `stop_reason`, and `usage`, or a +`ModelCallFailure`; check which before reading the response's fields. In `usage`, the counts of +cached and reasoning tokens are `None` when the provider doesn't report them. + +## Sampling settings + +`SamplingParams` has `max_tokens`, `temperature`, `top_p`, `stop_sequences`, `reasoning`, and a +provider-specific `extra`. `ReasoningParams` has `enabled`, `budget_tokens`, and `effort` +(`"minimal"` to `"max"`, clamped to the levels a provider offers). A setting left `None` leaves +the choice to the adapter or the provider, and which settings take effect depends on the adapter +and the model: `OpenAIResponsesAPI`, for example, doesn't apply `stop_sequences`. Leave a setting +unset unless your adapter and model support it, and don't expect `extra` settings to carry over +from one adapter to another. Write numeric fields with their declared types, such as +`temperature=0.0` rather than `0`, since stored records are checked strictly. + +`SamplingOverrides(inner=api, reasoning=...)` wraps an adapter and pins the reasoning settings of +every request; despite its name, it overrides reasoning only. + +Provider prompt-cache hints (`CacheableContentText`, `CacheableToolParam`, `CacheControl`) ask the +provider to cache a prompt's prefix. They are unrelated to fxtr's +[step cache](../../../fxtr/references/concepts/managing-the-step-cache.mdx), which keeps completed step results. diff --git a/plugins/fxtr/skills/fxtr/SKILL.md b/plugins/fxtr/skills/fxtr/SKILL.md new file mode 100644 index 0000000..15c3ef3 --- /dev/null +++ b/plugins/fxtr/skills/fxtr/SKILL.md @@ -0,0 +1,134 @@ +--- +name: fxtr +description: Set up, write, run, and view fxtr experiments, including new projects (`fxtr new`), steps and workflows, sweeps over models, prompts, or datasets, reductions, querying arrays (filter, aggregate, join, SQL), input datasets, reports, launching and resuming jobs, cached step results and reruns, and viewer renderers, results overviews, and the charts and visualizations in them. Use whenever starting or working in a fxtr project, creating or changing an experiment, or when a run reused results it should have recomputed (or recomputed results it should have reused). +--- + +# fxtr + +fxtr runs experiments as durable, inspectable DAGs. An experiment is Python code in a fxtr +project: steps do the computation, and workflows schedule steps over arrays with named +dimensions. A launch runs a workflow as a job on the project's Postgres database, which caches +every step's result, and the experiment viewer shows the job's DAG and data through the +project's renderers. + +The pages below are fxtr's documentation, grouped by the part of the work they cover. Read the +ones a task needs before writing code. [Core concepts](references/concepts/overview.mdx) +shows how the parts fit together, and its +[glossary](references/concepts/overview.mdx#glossary) defines the terms the pages use. + +## Where to read + +### Setting up a project + +Installing fxtr, and a project's configuration, dependencies, and database. + +| Task | Read | +| ----------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------- | +| Start a project, or change its configuration, dependencies, or database | [Installation](references/installation.mdx), [Project configuration](references/reference/project-configuration.mdx) | + +### Writing experiments + +Steps and workflows, the arrays they compute over, and the records they store. A new experiment +usually means reading Workflows and steps and Arrays and parallel computations. + +| Task | Read | +| ------------------------------------------------------------------------------------------ | -------------------------------------------------------------------------------------------- | +| See a small experiment written, launched, and rerun end to end | [Your first experiment](references/first-experiment.mdx) | +| Write or change steps and workflows: signatures, determinism, hermeticity, failures | [Workflows and steps](references/concepts/workflows-and-steps.mdx) | +| Shape data and sweeps: arrays, handles, mapping, alignment, reductions, replicas | [Arrays and parallel computations](references/concepts/arrays-and-parallel-computations.mdx) | +| Filter, select, aggregate, or join arrays, or write SQL over them, in a workflow or a step | [Array queries](references/reference/array-queries.mdx) | +| Define records: entities or structs, configurations, verdicts, reports | [Entities and custom types](references/concepts/entities-and-custom-types.mdx) | +| Import external datasets, pass configuration, and store entities | [Importing external datasets](references/guides/importing-external-datasets.mdx) | +| Call a language model, hold a conversation, or judge outputs | the adjacent [behaviors skill](../behaviors/SKILL.md) | + +### Visualizing experiments + +Renderers, which draw the project's entities, steps, and results in the viewer. Read Designing +views before drawing anything in a renderer. + +| Task | Read | +| -------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------- | +| Write and register renderers for entity types, steps, and workflows, including results overviews | [Visualization and custom renderers](references/concepts/visualization-and-custom-renderers.mdx), [Renderers](references/reference/renderers.mdx) | +| Design a view, chart, or visualization to fit the viewer | [Designing views](references/guides/designing-views.mdx) | +| Build the renderer package, fix a bundle that doesn't build or load, or add a package to a project | [The renderer package](references/reference/renderers.mdx#the-renderer-package) | + +### Running jobs + +Launching and following jobs, and the step cache that decides what a launch recomputes. Read +Running and viewing jobs before the first launch. + +| Task | Read | +| ------------------------------------------------------------------------------------------------------ | -------------------------------------------------------------------------- | +| Launch a job, watch it in the viewer, return a report and read results, resume, inspect, or share jobs | [Running and viewing jobs](references/guides/running-and-viewing-jobs.mdx) | +| Reuse or recompute cached results, resolve cache conflicts, change a step's logic | [Managing the step cache](references/concepts/managing-the-step-cache.mdx) | +| Look up a command's options | [CLI reference](references/reference/cli.mdx) | + +## Best practices + +Do these by default, unless the user asks otherwise. + +### Writing experiments + +* **Name every step and workflow.** Pass `name=` to each `@step` and `@workflow`. The step cache, + resumed jobs, and the viewer's renderers find a function by its + [registered name](references/concepts/workflows-and-steps.mdx#registered-names), and the + default name, taken from the function's module, changes when the function moves. +* **Write each docstring's first sentence for the viewer.** The operation's card shows it, so make + it a short imperative sentence saying what the operation does + ([Watch jobs in the viewer](references/guides/running-and-viewing-jobs.mdx#watch-jobs-in-the-viewer)). +* **Return a report from the root workflow**: one entity that references every result worth + presenting, for the launcher and the results overview to read + ([Read results](references/guides/running-and-viewing-jobs.mdx#read-results)). +* **Keep the graph's edges.** If an unavoidable reconstruction loses an edge in the viewer (an + array observed, transformed in Python, and added back as a source), tell the user. +* **Keep stale results out of a changed step.** The cache matches a step by its name and inputs, + not its code. After changing what a step computes, handle the change as + [Handling changes to step logic](references/concepts/managing-the-step-cache.mdx#handling-changes-to-step-logic) + describes before relaunching, and tell the user which results the next launch recomputes. + +### Visualizing experiments + +Write renderers in the project's renderer package, `views/`. If the project has none, ask the user +before [adding one](references/reference/renderers.mdx#adding-a-package-to-a-project). + +* **Give every entity type you introduce its renderers**, registered under its type tag: an + `EntityPanel` that leads with what matters in the entity, and an `EntityInlineLink` that names + it in a few words, since the default link shows only its type and a short ID. For a type with + many fields, add an `EntityHoverPreview` showing the few that identify it. Records from + behaviors already have renderers; the [behaviors skill](../behaviors/SKILL.md) says how to + include them. +* **Give every step you write an `InvocationPanel`**, registered under the step's name. It + becomes the step's default view, in its pane and in its card's preview, and it is handed every + invocation the viewer's selection picks out, up to all of a mapped step's. Show one in full and + several together (a table with a row per key, a chart, or a summary), loading their outputs + with `LoadEach` ([Step panels](references/reference/renderers.mdx#step-panels)). + Wrapping a view of one invocation in `singleInvocation` asks the reader to select one when + there are several, so keep it for the root workflow and for steps that are neither mapped nor + replicated. +* **Give the root workflow a results overview**: an `InvocationPanel` that reads the workflow's + report and presents the experiment's findings + ([Results overviews](references/concepts/visualization-and-custom-renderers.mdx#results-overviews)). +* **Check a view without opening it.** After changing a view, confirm it builds and run the + searches in [Check it](references/guides/designing-views.mdx#check-it), then tell the + user it is ready to look at in their viewer. Don't open the viewer in a browser yourself, or set + up Playwright or a headless browser to screenshot it, unless the user asks you to. + +### Running jobs + +* **Start the viewer before a launch.** When the user asks for a job to run, first start + `uv run fxtr view` in the background, as a long-running process of its own that you don't wait + on. When the project's viewer already runs, this only opens it, so it is safe every time. Each + launch then opens its job in a new tab of the viewer: tell the user it is there, and report the + job's ID and status from the launch's output. A launch that finds no viewer prints + `no viewer is running for this project` with the command to start one: start it then. +* **Keep the viewer running** across launches and renderer edits. It rebuilds the renderers when + their sources change, and the open viewer switches to them in place, so never ask the user to + reload or restart it. + +### Reading the documentation + +* **Know which documentation you have.** Distributed in the fxtr plugin for Claude Code or + Codex, this skill carries a copy of the pages, and `provenance.json` beside this file records + the fxtr and behaviors versions they describe. If the project depends on another fxtr version + (its `uv.lock` says which), tell the user. The plugin has its own version and update cycle; docs.transluce.ai has + the newest pages. diff --git a/plugins/fxtr/skills/fxtr/provenance.json b/plugins/fxtr/skills/fxtr/provenance.json new file mode 100644 index 0000000..24d2cd0 --- /dev/null +++ b/plugins/fxtr/skills/fxtr/provenance.json @@ -0,0 +1,11 @@ +{ + "skill": "fxtr", + "documentation": "https://docs.transluce.ai/fxtr", + "packages": { + "fxtr": "0.0.1a0", + "behaviors": "0.0.1a0" + }, + "revision": "216737e51a92f5a0a83f7629af73d5a733fceb19", + "modified": false, + "content_sha256": "c579b0bc6dbf02f109faa139ff0bee821f5f4010d549f53ec8990b8997fefdf8" +} diff --git a/plugins/fxtr/skills/fxtr/references/INDEX.md b/plugins/fxtr/skills/fxtr/references/INDEX.md new file mode 100644 index 0000000..74de3e1 --- /dev/null +++ b/plugins/fxtr/skills/fxtr/references/INDEX.md @@ -0,0 +1,36 @@ +# The fxtr documentation + +The pages of https://docs.transluce.ai/fxtr, copied with this skill: `provenance.json`, +beside its `SKILL.md`, says from which version. Each is listed with what it covers, +in the order of the site's navigation. + +## Get started + +- [What is fxtr?](index.mdx): fxtr is an experiment orchestration system for verifiable AI-powered analyses +- [Installation](installation.mdx): Install fxtr and get started with an example project +- [Your first experiment](first-experiment.mdx): Write a small experiment with a mapped step and a reduction, launch it, and see how caching works. + +## Concepts + +- [Core concepts](concepts/overview.mdx): How projects, steps, workflows, jobs, arrays, and entities fit together. +- [Workflows and steps](concepts/workflows-and-steps.mdx): How an experiment is written: workflows describe the structure, steps do the work. +- [Arrays and parallel computations](concepts/arrays-and-parallel-computations.mdx): The data model: arrays indexed by named dimensions, and how steps are mapped over them. +- [Entities and custom types](concepts/entities-and-custom-types.mdx): Stored records with identities of their own: when to define one, how to store and load it, and how it differs from a struct in an array. +- [Managing the step cache](concepts/managing-the-step-cache.mdx): How fxtr reuses step results across jobs, when it refuses to, and how to control it. +- [Visualization and custom renderers](concepts/visualization-and-custom-renderers.mdx): How the viewer draws a job, and how renderers teach it to show your own types, steps, and results. + +## Guides + +- [Running and viewing jobs](guides/running-and-viewing-jobs.mdx): Launch jobs from Python or the command line, watch them in the viewer, read their results, and resume, inspect, or share them. +- [Importing external datasets](guides/importing-external-datasets.mdx): Bring external datasets into fxtr as typed arrays for processing in workflows. + +## Reference + +- [CLI reference](reference/cli.mdx): The fxtr command-line interface. + +## Other pages + +- [Designing views](guides/designing-views.mdx): How a renderer's views should look and behave in the viewer: surfaces, type, color, charts, links, controls, and how to check them. +- [Array queries](reference/array-queries.mdx): Filter, reshape, aggregate, and join arrays with the query builder or SQL, in a workflow or in a step. +- [Project configuration](reference/project-configuration.mdx): What fxtr new writes, and the settings in pyproject.toml and fxtr.local.toml. +- [Renderers](reference/renderers.mdx): The renderer API: slots and registration, loading entities and invocations, step panels, links, panel state, the kit, and building the renderer package. diff --git a/plugins/fxtr/skills/fxtr/references/concepts/arrays-and-parallel-computations.mdx b/plugins/fxtr/skills/fxtr/references/concepts/arrays-and-parallel-computations.mdx new file mode 100644 index 0000000..b5d623b --- /dev/null +++ b/plugins/fxtr/skills/fxtr/references/concepts/arrays-and-parallel-computations.mdx @@ -0,0 +1,486 @@ +--- +title: "Arrays and parallel computations" +description: "The data model: arrays indexed by named dimensions, and how steps are mapped over them." +--- + +An **array** is a collection of values indexed by named **dimensions**. Every input and output of +a step or workflow is an array, and a single value is simply an array with no dimensions. Arrays +are how fxtr represents regular structures of data, and how it expresses parallel work: mapping a +step over an array's dimensions runs it once per key and stacks the results back into a new array. + +If you have used xarray or a labeled dataframe, the ideas will feel familiar. If you haven't, the +picture to hold is a table whose rows are addressed by names rather than positions: + +| model | prompt | value | +| ------------ | -------- | ------ | +| `"baseline"` | `"math"` | `0.72` | +| `"baseline"` | `"code"` | `0.64` | +| `"variant"` | `"math"` | `0.81` | + +This array has two dimensions, `model` and `prompt`, both with string keys, and `float` values. +Its type is written `[model: str, prompt: str]: float`. The `variant` model has no `code` score, +and that is fine: arrays may be sparse. What they may not be is irregular. Every value has the same +shape, every key appears at most once, and there is one value per key. + +## Structure of an array + +### Dimensions + +Each dimension has a name and a key type, either `int` or `str`. A row's key is the tuple of its +values along every dimension, in the array's dimension order, and no two rows share a key. Keys +don't need to form a full grid, so you can have different numbers of examples per model without +padding. + +Dimension *names* are what fxtr matches on. When two arrays both have a `model` dimension, fxtr +treats them as the same axis and lines them up by key. Dimension *order* is only presentation: it +fixes how key tuples are written, and nothing else. Keep names specific and consistent across an +experiment, since the same name on two unrelated axes will make fxtr try to align them. + +### Values + +Every value in an array follows one **value schema**. A schema is one of: + +* a primitive: `int`, `float`, `bool`, `str`, or `bytes`; +* a reference to an **[entity](entities-and-custom-types.mdx)**, a stored record with its own ID; +* a list of values of one schema; +* a struct with a fixed set of named fields, each with its own schema; +* a nullable version of any of the above, written `X | None`. + +A struct-valued array looks a lot like a dataframe: the dimensions are the index, and the struct's +fields are the columns. The analogy is useful, but fxtr treats the dimensions and fields as +fundamentally different things. Dimensions address rows, must be unique for each array, and are what +steps are mapped over and what arrays are aligned on. Struct fields are part of the value; they do not +have to be unique, are not mapped over, and can be manipulated via operations like `field(...)` to +extract individual parts. (Similarly, lists inside a value are not mapped over, but can be read by +steps or SQL queries; see [Turning lists into dimensions](#turning-lists-into-dimensions).) + +Regularity is enforced. Values are checked against the schema when an array is built, and a +value that doesn't fit is an error. That keeps array operations and queries simple and fast, but it +means arrays cannot hold data whose shape varies from row to row. If a row needs a tagged union, a +blob of JSON, or an object with optional parts, make that object an entity and put a reference to +it in the array. [Entities and custom types](entities-and-custom-types.mdx) explains +when to do which. + +### Writing types in Python + +An array type is a pair of the dimensions, written as a string, and the value schema, written as +a Python type: + +```python +from typing import TypedDict + +from fxtr.core.entities import BoundID + + +class Case(TypedDict): + prompt: str + expected: str + + +("[prompt: str]", str) # strings indexed by prompt +("[model: str, sample: int]", float) # floats indexed by model and sample +("[case_id: str]", Case) # structs with two string fields +("[doc: str]", BoundID[Document]) # references to Document entities +("[]", int) # a single int +``` + +A `TypedDict` describes a struct, `BoundID[T]` describes a reference to an entity of type `T`, +and `X | None` makes a value nullable. Dataclasses and other unions are not value schemas; they +belong in entities. + + + The single-string form `"[prompt: str]: str"` is how fxtr *displays* a type. The constructors + take the pair `("[prompt: str]", str)` and reject the display form. + + +## Creating arrays + +`Array` lives in `fxtr.entity_defns.array`. The constructors take the data and the type: + +```python +from fxtr.entity_defns.array import Array + +prompts = Array.from_items( + [(("greeting",), "Hello there"), (("question",), "Why is the sky blue?")], + ("[prompt: str]", str), +) +cases = Array.from_records( + [{"case_id": "a1", "prompt": "2 + 2", "expected": "4"}], + ("[case_id: str]", Case), +) +budget = Array.scalar(200, int) + +assert str(prompts.type) == "[prompt: str]: str" +assert str(cases.type) == "[case_id: str]: struct{expected: str, prompt: str}" +assert str(budget.type) == "[]: int" +``` + +* `Array.from_items` takes `(key, value)` pairs. A key is a tuple in dimension order, or a + dictionary from dimension name to key. +* `Array.from_records` builds a struct-valued array from flat records: each record holds the + dimension keys and the struct's fields together, the way a CSV row does. +* `Array.scalar` builds a single value. +* `infer_from_items`, `infer_from_records`, and `infer_scalar` work the type out from the data when + that is unambiguous. An empty array always needs an explicit type. + +Reading is equally direct: `array[("greeting",)]` gives one value, `array.items()` yields +`(key, value)` pairs in key order, `array.values()` yields the values, and `array.item()` gives the +one value of a scalar. Arrays are immutable once built. + +### Getting data into an experiment + +A workflow receives its inputs from whoever launches the job, and that launcher is the one place in +a fxtr project that may read files from disk (see +[Hermeticity](workflows-and-steps.mdx#hermeticity-and-handling-external-state)). There +are two recommended ways to get an external dataset in: + +* **Construct the array in the launcher.** Read the file, validate it, build a typed array, and + pass it to `run_job`. fxtr stores the array with the job, so the job keeps the exact data it ran + on even if the file changes later. +* **Persist it as a project array.** A **project array** is a named, typed, editable array kept + in the project's database. You fill it with the project client, from a script or notebook, and + each launch takes an immutable **snapshot** to pass as an input. This suits a dataset you curate + over time and launch many jobs against. + +```python +await client.create_array("eval_cases", ("[case_id: str]", Case)) +await client.replace_array("eval_cases", cases) +snapshot = await client.snapshot_array("eval_cases") +await run_job(client, experiment, {"cases": snapshot.array}, root="eval-v1") +``` + +Either way, key rows by stable IDs from the source rather than by row number, so that adding or +reordering rows doesn't change which cached result belongs to which row. Configuration that is +fixed in the experiment's code can go in with `context.add_source_array(...)` inside the workflow. + +## Array versus ArrayHandle + +Two Python types refer to arrays, depending whether the value is currently known: + +* An **`Array`** contains concrete data. Arrays can be constructed in launcher code and are used + inside steps, where every input arrives as an `Array` (or is unwrapped to a bare value, for a + scalar). +* An **`ArrayHandle`** is a lazy reference to an array that will exist later, such as a step's + future result. It appears inside workflows, where inputs arrive as handles and every scheduling + operation returns a new handle. + +A handle knows its type immediately. That is what lets a workflow pass handles from one step to +the next, building up a graph of dependencies, before anything has run. When a workflow does need +the data in order to do [data-dependent control flow](workflows-and-steps.mdx#data-dependent-control-flow), +`await handle.observe()` waits for the array and `await handle.item()` waits for a scalar's value. +(You should generally avoid using these just to transform data inside a workflow body; prefer to use +a step or SQL query instead.) + +`handle.expect_type(("[model: str]", float))` checks a handle's type where you build it and raises +if your assumption is wrong, which is a cheap way to make a workflow's shapes readable. + +## Mapping logic over arrays + +Scheduling a step is one operation, `run_step`, and the same operation maps it over any number of +dimensions. `run_child_workflow` follows exactly the same rule. The `map_over` argument lists the +dimensions to map over, and the runner: + +1. joins the inputs' keys along those dimensions, broadcasting any input that lacks one of them; +2. calls the function once per joint key, passing each input's slice with the mapped dimensions + removed; +3. stacks the results into one array, whose dimensions are the mapped ones followed by whatever the + function itself returns. + +```mermaid +flowchart LR + M["models
[model]"] --> R["run_step(answer, map_over=[model, prompt])"] + P["prompts
[prompt]"] --> R + C["judge config
[] (broadcast)"] --> R + R --> O["answers
[model, prompt]"] +``` + +Dimensions not listed in `map_over` are **core dimensions**: they stay inside each call, and the +function receives them whole. With `map_over` left empty, which is the default, the function is +called once with the entire input arrays. Mapping is never implicit. With +[replicas](#replicas), the replica dimensions come between the mapped dimensions and the +function's own, and the function's own dimensions can't reuse a mapped or replica dimension's +name. + +### Start from one call + +The reliable way to design a mapped step is to write its signature for **one call**, then decide +which input dimensions should produce multiple calls. Suppose each call of `run_task` takes one +model, one task, and all the few-shot examples for its model: + +| Parameter | One call expects | Supplied array | +| ---------- | ---------------- | ----------------------- | +| `model` | a single `str` | `[model]: str` | +| `task` | a single `str` | `[task]: str` | +| `examples` | `[example]: str` | `[model, example]: str` | + +```python +# models: [model: str]: str +# tasks: [task: str]: str +# examples: [model: str, example: int]: str +results = context.run_step( + "run_tasks", + run_task, + {"model": models, "task": tasks, "examples": examples}, + map_over=["model", "task"], +) +# One call per (model, task); each call sees its model's [example] slice. +# If run_task returns a float, the result is indexed by the mapped dimensions: +results.expect_type(("[model: str, task: str]", float)) +``` + +In this call: + +* `model` and `task` have no dimensions in common, so they cross: two models and three tasks make + six calls. +* `examples` shares the `model` dimension, so each call receives the `[example]` slice for + its own model (with this same slice used for all tasks). Different models may have different numbers of examples, which are then given directly to the step function. +* The `[model, task]` dimensions are prepended to the dimensions of the step's result. (If the step returns a scalar, the result will have just these two dimensions.) + +It's important to distinguish between **mapping** a step over inputs and **passing the whole +set** to one call. Mapping `count_words` over `prompt` gives one call per prompt and a result +indexed by prompt. Passing the whole `[prompt]` array to `mean_length` with no `map_over` gives one +call that sees every prompt and returns one number. A reduction is just a step whose signature +keeps the dimensions it reduces over as core dimensions: + +```python +@step(name="eval.mean_score") +async def mean_score( + context: StepContext, + scores: Annotated[Array[float], "[task: str, sample: int]"], +) -> float: + """Average a model's scores across tasks and samples.""" + values = list(scores.values()) + return sum(values) / len(values) + + +# scores: [model: str, task: str, sample: int]: float +per_model = context.run_step("mean_scores", mean_score, {"scores": scores}, map_over=["model"]) +# Mapping over model leaves [task, sample] as the core dimensions each call receives. +per_model.expect_type(("[model: str]", float)) +``` + +Each call receives one model's `[task, sample]` scores, and `per_model` has type `[model]: float`. + +### Alignment rules + +Inputs line up by dimension name and key: + +* Inputs that share a dimension align by key. (If two independent axes happen to + share a name, rename one with `handle.rename({...}).collect()` first.) +* An input lacking a mapped dimension is broadcast along it. A scalar is broadcast to every call. +* Inputs with entirely disjoint dimensions become a cross product (since each input is broadcast across all of its missing dimensions). + +By default (with `join="exact"`), inputs fail to align if doing so would drop any input's key, so a missing row is an +error rather than a silent skip. Sparse arrays are still allowed as long as every input agrees on which keys it +has. You can choose to instead use an inner join by passing `join="inner"`, which only runs the step for keys that are present in all inputs (after broadcasting). + +### Selecting parts of an input + +`handle.field("name")` selects one field of a struct-valued array while keeping its dimensions. +This is the usual way to pass pieces of a dataset to a step that expects simple values: + +```python +# cases: [case_id: str]: Case (a struct with prompt and expected fields) +# judge_config: []: BoundID[JudgeConfig] (a scalar, broadcast to every call) +questions = cases.field("prompt").collect() +questions.expect_type(("[case_id: str]", str)) # the field, still keyed by case_id + +answers = context.run_step( + "answer", + answer, + {"question": questions, "config": judge_config}, + map_over=["case_id"], +) +answers.expect_type(("[case_id: str]", BoundID[Answer])) +``` + +Because the field keeps the dataset's `case_id` keys, the result lines up with the dataset and +with anything else derived from it. The same trick, combined with a query, is how you run a step +over only some of the inputs (see +[Running a step over a subset of the inputs](#running-a-step-over-a-subset-of-the-inputs)). +`collect()` is explained in +[Querying, filtering, and aggregating arrays](#querying-filtering-and-aggregating-arrays). + +### Replicas + +`replicas={"sample": range(3)}` repeats each call three times with the same inputs, under +distinct cache addresses, so the three calls are independent samples. The result gains a `sample` +dimension after the mapped ones. The step body does not see the replica number, so `replicas` usually +makes sense when the step is fundamentally nondeterministic. (If a step needs a random seed, pass seeds +as an ordinary input array and map over its dimension.) + +## Querying, filtering, and aggregating arrays + +Between steps you often need to filter, reshape, aggregate, or join arrays. fxtr has a **query +builder**, a set of methods such as `filter`, `select`, `aggregate`, and `join`, and an SQL entry +point, `fxtr.sql(...)`, for anything the builder doesn't express. Both describe a query lazily, and +`collect()` runs it. Queries are executed by DuckDB, so they are fast on large arrays. + +The same methods work on an `Array` and on an `ArrayHandle`, with one difference in what +`collect()` does: + +* **On arrays** (in a step, a launcher, or a notebook), `await query.collect()` runs the query now + and returns an `Array`. +* **On handles** (in a workflow), `query.collect()` is not awaited. It adds the query to the + workflow's graph as a **query node**, which the runner computes when its inputs are ready, and + returns a handle to it. The rows never pass through the workflow body. + +```python +from typing import TypedDict + +from fxtr.entity_defns.query import count, dim, value + + +class Run(TypedDict): + score: float + error: str | None + + +class Summary(TypedDict): + runs: int + mean_score: float | None # mean() of no rows is null, so the field is nullable + + +# In a workflow, with `runs` a handle of type [model: str, run: int]: Run. +# Keep the runs that finished, then summarize each model. Nothing runs here. +finished = runs.filter(value().field("error").is_null()).collect() +finished.expect_type(("[model: str, run: int]", Run)) # a filter keeps the dimensions + +per_model = finished.aggregate( + {"runs": count(), "mean_score": value().field("score").mean()}, + group_by={"model": dim("model")}, +).collect() +per_model.expect_type(("[model: str]", Summary)) # group_by names become the dimensions + +# In a step, with `runs` an Array of the same type: the same query, run now. +summary = await runs.filter(value().field("error").is_null()).aggregate( + {"runs": count(), "mean_score": value().field("score").mean()}, + group_by={"model": dim("model")}, +).collect() +assert str(summary.type) == "[model: str]: struct{mean_score: float?, runs: int}" +``` + +`dim("model")` reads a dimension key, `value()` reads the row's value, and `value().field("x")` +reads a field of a struct value. Predicates use `&`, `|`, `~`, `is_in`, and `is_null` rather than +Python's `and`, `or`, `not`, and `in`, and they follow SQL's rules about nulls. `fxtr.sql` takes a +DuckDB `SELECT` over named tables, each of which has a `dims` column and a `value` column. + +### Running a step over a subset of the inputs + +A common need is to run a step over some of a dataset rather than all of it: only the hard cases, +only the pairs of model and harness that make sense, only the rows that survived an earlier +filter. The pattern has three parts, and it works because queries and field selection both keep +the array's dimension keys: + +1. **Keep the inputs together in one struct-valued array.** Each row holds everything the step + will need for that key, as fields. +2. **Select the rows with a query.** Filter the handle, or build an array holding only the + allowed keys. The result has the same dimensions and fewer rows. +3. **Map the step over the selected rows, passing the fields it needs with `.field()`.** The + fields keep the selected keys, so the join produces exactly those calls. + +Suppose `Case` also has a `difficulty` field, and only the hard cases should be answered: + +```python +# cases: [case_id: str]: Case, with fields prompt, expected, and difficulty +hard = cases.filter(value().field("difficulty") == "hard").collect() +hard.expect_type(("[case_id: str]", Case)) # fewer rows, same dimensions + +answers = context.run_step( + "answer_hard", + answer, + {"question": hard.field("prompt").collect(), "config": judge_config}, + map_over=["case_id"], +) +answers.expect_type(("[case_id: str]", BoundID[Answer])) # one row per hard case +``` + +The same pattern restricts which **combinations** run. Passing a `[model]` array and a +`[harness]` array to a mapped step runs every pair. To run only some pairs, first build the full +grid as one struct-valued array with `context.combine`, then filter it. Here each model and each +harness records its provider, and only cross-provider pairs should run: + +```python +# models: [model: str]: ModelInfo (fields name and provider) +# harnesses: [harness: str]: HarnessInfo (fields name and provider) +pairs = context.combine({"model": models, "harness": harnesses}) +# combine broadcasts each input along the dimension it lacks, giving the full grid: +# [model: str, harness: str]: struct{model: ModelInfo, harness: HarnessInfo} + +allowed = pairs.filter( + value().field("model").field("provider") != value().field("harness").field("provider") +).collect() +# Same dimensions, only the rows where the predicate holds. + +results = context.run_step( + "run_allowed", + run_task, + { + "model": allowed.field("model").field("name").collect(), # [model, harness]: str + "harness": allowed.field("harness").field("name").collect(), # [model, harness]: str + "task": tasks, # [task]: str + }, + map_over=["model", "harness", "task"], +) +results.expect_type(("[model: str, harness: str, task: str]", float)) # allowed pairs × tasks +``` + +Both selected fields carry the joint `(model, harness)` keys of the filtered grid, so the join +produces the allowed pairs crossed with the tasks and nothing else. If another input, such as +per-model examples, has rows for models that end up unselected, pass `join="inner"` so those rows +are dropped rather than reported as missing. When the allowed pairs are an explicit list rather +than a rule, build them in the launcher with `Array.from_records` over the joint keys and pass +them in as an input; the selection step is the same from there. + +When the selection depends on earlier results, compute it with a query over those results' +handles, or in a step that takes them. Don't observe the results, choose rows in Python, and add +the choice back as a source array: the viewer would lose the edge from the results to the +selection. Because a query node isn't a step, selecting this way costs nothing in the cache, and +the query shows up in the viewer as part of the argument that reads it. + +### Turning lists into dimensions + +Sometimes a step returns a struct that contains a list. For example, an extraction step might read +each answer and return a summary together with the list of claims the answer makes: + +```python +class Extraction(TypedDict): + summary: str + claims: list[str] +``` + +The step's result has type `[answer: str]: Extraction`. The next step must check each claim on +its own, so it must be mapped over the claims, but the claims are inside a list, not a dimension. +We can solve this with a SQL query that selects the `claims` field and unnests it: `unnest` +gives one row per element, and `generate_subscripts` gives each element's position, which becomes +the new dimension's key: + +```python +# extractions: [answer: str]: Extraction +claims = fxtr.sql( + """ + SELECT {answer: dims.answer, i: generate_subscripts(value.claims, 1) - 1} AS dims, + unnest(value.claims) AS value + FROM extractions + """, + tables={"extractions": extractions}, + output_type=("[answer: str, i: int]", str), +).collect() +claims.expect_type(("[answer: str, i: int]", str)) # one row per claim +``` + +Now a step can be mapped over `answer` and `i`, and the `answer` dimension connects each claim back to +its extraction. To fold a dimension back into a list, you can use `list(value ORDER BY dims.i)` with a +`GROUP BY` on the dimensions that remain. (The query builder has no `unnest`, so this is one of +the places where you need SQL.) + + + Query nodes are not cached: the runner recomputes them whenever a job resumes and in every + later job. That is harmless for deterministic queries. It is not for SQL that asks for varying + values, such as `random()` or `now()`, and it can be a problem for summing floating-point values + whose order of summation is not fixed. Put random draws, timestamps, and sensitive floating-point + aggregations in steps, whose results are cached, rather than in workflow queries. + + +The full set of operations, the places where queries differ from Python, and more worked +examples are in the [array query reference](../reference/array-queries.mdx). diff --git a/plugins/fxtr/skills/fxtr/references/concepts/entities-and-custom-types.mdx b/plugins/fxtr/skills/fxtr/references/concepts/entities-and-custom-types.mdx new file mode 100644 index 0000000..97b6819 --- /dev/null +++ b/plugins/fxtr/skills/fxtr/references/concepts/entities-and-custom-types.mdx @@ -0,0 +1,204 @@ +--- +title: "Entities and custom types" +description: "Stored records with identities of their own: when to define one, how to store and load it, and how it differs from a struct in an array." +--- + +An **entity** is a stored record with a type and an **ID derived from its content**. Where an +array holds many regular values addressed by key, an entity is a single thing you can point at. +Arrays hold references to entities, entities can hold references to other entities, and the viewer +knows how to follow every reference and show what it finds. + +Entities are for three kinds of data: + +* **Large objects you want to refer to later as a whole.** A conversation transcript, a document, + a generated report. You pass references around and load the content only where it's needed. +* **Objects with irregular structure.** Array values must follow one fixed schema, but an entity + can hold a tagged union, nested optional parts, or a blob of JSON, because it is serialized as a + unit rather than laid out in columns. +* **Things other objects should point to.** A judge's verdict about a transcript naturally + references the transcript. A result should reference the configuration that produced it. Entity + references make these links explicit, and the viewer renders them as links you can click. + +Because an entity's ID is a hash of its content, identical content always has the same ID and is +stored once. Storing the same model configuration from a hundred steps costs one record, and two +jobs that reference it reference the very same thing. + +## Defining entities + +Most entity types are frozen dataclasses that inherit from `DataclassEntity`: + +```python +from dataclasses import dataclass + +from fxtr.core.entities import BoundID +from fxtr.entity_defns.dataclass_entity import DataclassEntity + + +@dataclass(frozen=True) +class ModelConfig(DataclassEntity, fxtr_type="org.example.qa.ModelConfig.v1"): + model_name: str + temperature: float + + +@dataclass(frozen=True) +class Answer(DataclassEntity, fxtr_type="org.example.qa.Answer.v1"): + config: BoundID[ModelConfig] + question: str + text: str +``` + +The `fxtr_type` keyword is the entity's **type tag**. Every entity type must have one, it must be +unique, and it is how the system refers to the type everywhere: in storage, in array schemas, and +in the viewer, where renderers are keyed by it. Use a reverse-domain name, and include a version so +that a changed type can be introduced under a new name while old records keep theirs. + +Fields may be primitives, tuples and frozen sets, nested dataclasses, `BoundID[T]` references to +other entities, nullable and union types, and recursive JSON values. The `Answer` above points at +its `ModelConfig` by reference, so every answer records exactly which configuration produced it +without copying the configuration into every record. + +### Inline tagged unions + +A field whose value can be one of several shapes, say a step outcome that is either a score or a +refusal, is written as a union of **tagged dataclasses**. A `TaggedDataclass` carries a type tag +like an entity does, but it is not stored on its own: it is serialized inline, inside whatever +entity contains it, and the tag tells the loader which variant it is reading. + +```python +from fxtr.entity_defns.dataclass_bases import TaggedDataclass + + +@dataclass(frozen=True) +class Scored(TaggedDataclass, fxtr_type="org.example.qa.Scored.v1"): + score: float + + +@dataclass(frozen=True) +class Refused(TaggedDataclass, fxtr_type="org.example.qa.Refused.v1"): + reason: str + + +@dataclass(frozen=True) +class Verdict(DataclassEntity, fxtr_type="org.example.qa.Verdict.v1"): + answer: BoundID[Answer] + outcome: Scored | Refused +``` + +This is the kind of structure an array value cannot hold, and one of the main reasons to reach for +an entity. + + + fxtr currently converts dataclasses to and from their stored form with the + [cattrs](https://catt.rs/) library, with strict settings: unknown fields, missing fields, and + values of the wrong type are rejected rather than coerced. This implementation detail may change + in a future version; the dataclass surface described here is what to rely on. + + +### Advanced use: low-level entity classes + +Underneath the dataclass layer, an entity is defined by one ability: it can serialize itself to a +canonical form, and something can deserialize that form back into an object. The canonical form is +a tree of plain values (`None`, numbers, booleans, strings, bytes, lists, string-keyed mappings, +and entity IDs) with a `$type` key holding the type tag, encoded as DAG-CBOR and hashed to produce +the ID. Any class can be an entity by subclassing `Entity` and implementing its abstract methods: + +```python +from fxtr.core.entities import Entity, StaticTypeLoader + + +class PointLoader(StaticTypeLoader["Point"]): + _static_type_id = "org.example.geometry.Point.v1" + + def convert(self, data): + return Point(data["x"], data["y"]) + + +class Point(Entity): + def __init__(self, x: float, y: float): + self.x, self.y = x, y + + @classmethod + def fxtr_type_id(cls): + return "org.example.geometry.Point.v1" + + def _fxtr_serialize(self): + return {"$type": "org.example.geometry.Point.v1", "x": self.x, "y": self.y} + + def _fxtr_loader(self): + return PointLoader() +``` + +`_fxtr_serialize` produces the canonical form and `_fxtr_loader` names the **loader** that turns +it back into an object. A loader's `convert` receives the stored data and returns the value; +`StaticTypeLoader` is the base for a loader that handles exactly one type tag. + +Note that the low-level protocol is usually unnecessary; prefer to use `DataclassEntity` if possible. + +## Storing and loading entities + +Storing an entity returns a `BoundID`: the entity's ID together with the loader that knows how to +read it back. You can store from a step context, a workflow context, or the project client in a +launcher, and all three expose the same two methods: + +```python +ref = await context.store(ModelConfig(model_name="gpt-6-luna", temperature=0.2)) +config = await context.load(ref) +``` + +A `BoundID` is cheap to pass around and to put in other entities and arrays. A bare `EntityID`, +such as one read out of an array of untyped IDs, needs its type to be loaded: +`await context.load(entity_id, as_type=ModelConfig)`. + +The only way to get an entity back is to have its ID, and storing an entity gives it no name and +puts it in no list you can browse. So for an entity to be retrievable later, something that is +itself persisted in the database must reference it: a job's inputs, which the job records; a +job's results, which hold the references of the entities steps return; or a project array. An +entity that nothing references can't be found again, so when you store entities in a launcher, +pass the references on as a job input or write them into a project array. + +### Entities in arrays + +Arrays hold references, never entities themselves. The common pattern is to store each entity and +then build the array from the references, with `BoundID[T]` as the value type: + +```python +refs = [((name,), await client.store(config)) for name, config in configs.items()] +models = Array.from_items(refs, ("[model: str]", BoundID[ModelConfig])) +``` + +Two shorthands cover the step side. A step annotated to **receive** an entity, as in +`config: ModelConfig`, gets the loaded object rather than the reference. A step annotated to +**return** an entity type, as in `-> Answer`, may return the entity itself: fxtr stores it and +records the reference as the step's result, so the mapped results of such a step form an array of +`BoundID[Answer]` without any explicit store calls. + +## Entities versus array structs + +There are two ways to put structured data in an array: a **struct value**, written as a +`TypedDict`, whose fields live inside the array as columns, or an **entity**, stored separately +and referenced by ID. The rule of thumb: + +* Use a **struct** when the data is regular and exists to organize values for this step or + workflow, without needing an identity of its own. "A collection of an X and a Y", dataset rows, + sweep coordinates, numbers extracted for aggregation, summary statistics. +* Use an **entity** when you want to refer to the thing by ID and point at it from elsewhere, when + it is naturally a single thing rather than a collection of values (a transcript, a verdict, a + configuration, a report), or when its shape is irregular. + +The practical differences follow from that: + +| | Struct value | Entity | +| -------------- | --------------------------------------------------- | -------------------------------------------------- | +| Identity | None of its own; lives inside its array | Its own ID; identical content stored once | +| In a workflow | Fields visible: `handle.field("score")` selects one | Opaque: a step that takes it depends on all of it | +| In the viewer | Shown as columns of the array | Linked, and shown by a renderer chosen by its type | +| Allowed shapes | Primitives, lists, nested `TypedDict`s, nullables | Any dataclass fields, including unions | + +A step can combine the two: store an entity for the record that deserves one, and return a struct +row holding the fields you'll aggregate next to a `BoundID` pointing at the entity. The aggregation +then reads columns, and anyone who wants the detail follows the reference. + + + Coming soon: We plan to make it possible to efficiently extract parts of entities into an array struct, to + make it possible to efficiently process and summarize individual fields of entity data. + diff --git a/plugins/fxtr/skills/fxtr/references/concepts/managing-the-step-cache.mdx b/plugins/fxtr/skills/fxtr/references/concepts/managing-the-step-cache.mdx new file mode 100644 index 0000000..b3c3ec3 --- /dev/null +++ b/plugins/fxtr/skills/fxtr/references/concepts/managing-the-step-cache.mdx @@ -0,0 +1,255 @@ +--- +title: "Managing the step cache" +description: "How fxtr reuses step results across jobs, when it refuses to, and how to control it." +--- + +Every step's result is stored in the project's **step cache**. A later job that asks for the +same step on the same inputs reuses the result instead of running the step again, which +makes a fxtr experiment cheap to iterate on. You can edit the workflow to widen a sweep, add a +new analysis step, or fix a bug downstream, then launch a new job and pay only for the work that is +actually new. Your experiment code never has to manage data versions or check whether a result +already exists; it describes what should be computed, and fxtr uses the cache to determine what +values you already have. + +The cache is also where fxtr is deliberately conservative. By default, if a step requested by a job +doesn't match the cached result, the job is *interrupted* and the conflict must be manually cleared, +in case this change was unintentional. This page describes how to understand these cache conflicts +and how to clear them. + +## Mental model + +Four facts describe the cache: + +* **Every step invocation has a cache address.** Within one job, a cache address always resolves + to exactly one value. By default each invocation a workflow schedules gets a unique address built + from its position in the job, so you rarely need to think about addresses unless you want more control + over how results are shared between jobs. +* **Every invocation also has a fingerprint:** the step's registered name together with the exact + input data it received. Two invocations with the same fingerprint would compute the same thing + (or, for a nondeterministic step, a sample of the same thing). +* **There is one shared, mutable cache for the whole project**, keyed by address. Every job in + the project's database schema reads and writes it. Separately, **each job keeps its own + immutable record** of which result it took at each address. Clearing or overwriting a cache + entry later never changes what a finished job used. +* **A lookup compares fingerprints.** When a job reaches a step, it looks at the entry at the + step's address: + +```mermaid +flowchart TD + A["job reaches a step at address K"] --> B{"entry at K?"} + B -- "none" --> R["run the step, store the result at K"] + B -- "same fingerprint" --> U["reuse the result
(wait for it if still running)"] + B -- "different fingerprint" --> C["cache conflict:
step does not run, job stops"] +``` + +A **cache conflict** means the address already holds a result computed from different inputs, or +by a step with a different name. Reusing it would give a wrong answer; replacing it would destroy +work some other job may depend on. So by default the job stops once nothing else can proceed, and +reports the conflicting addresses with the commands that resolve them. Nothing is lost: every step +that could run has run, and the job can be resumed once you have decided what to do. + +A failed or abandoned entry (from a step that raised or a process that died) is not a conflict. +The next job to ask for that address simply runs the step. + +## Cache addresses + +Every operation in a job has a **pathname**, a readable address built from the names you give +things. The root workflow's pathname is the `root` you pass when launching the job, `"root"` by +default. Each step, child workflow, and source array adds its local name under its parent, which +must be unique within that parent, and a mapped or replicated invocation adds its keys, written as +JSON values in dimension order: + +| Pathname | Refers to | +| ------------------------------------------- | --------------------------------------------------------------------- | +| `qa-v1` | The root workflow of a job launched with `root="qa-v1"` | +| `qa-v1/answer[model:"small",prompt:"math"]` | One invocation of the `answer` step, mapped over `model` and `prompt` | +| `qa-v1/summarize[model:"small"]/mean` | An unmapped step inside a mapped child workflow | +| `qa-v1/sample[prompt:"math",sample:2]` | A step invocation with a replica key | + +A step's pathname is its default cache address. That has two consequences worth internalizing: + +* **Two jobs share results when their pathnames match.** The same `root`, the same local names + along the way, and the same mapped keys give the same address. Rename a step in the workflow, + move it into a child workflow, or change a dimension's name, and its results get new addresses, + so the next job computes them afresh. +* **The root name is a namespace.** Because every address starts with `root`, keeping the root the + same across launches is how you reuse results, and changing it is how you start from a clean + slate without touching anything. A new root is the simplest way to draw fresh samples from a + nondeterministic step, since the job ID is not part of the address and a new job under the old + root would reuse the old samples. To draw fresh samples on every launch, generate the root in + the launcher, as in `root=f"qa-{uuid.uuid4()}"`, and never inside a workflow. + +Names are not escaped, so never put `/`, `[`, or `]` in a local name, and make names specific: +`comparison_judge_config` reads better in an address than `config`. + +### Customizing an address + +Occasionally you want a step's results to be shared under an address that does not follow from +its position, for example a slow preprocessing step whose results several experiments should +share whatever their roots are. `run_step` and `run_child_workflow` take an +`override_cache_address_prefix` argument that replaces the pathname-derived prefix. Mapped and replica +keys are still appended to it, and a child workflow passes its prefix down to everything beneath +it. + +```python +clusters = context.run_step( + "cluster_transcripts", + cluster_transcripts, + {"transcripts": transcripts}, + override_cache_address_prefix="shared/clusters-v1", +) +``` + +Within a single job, an address must resolve to a single value. If two invocations in one job end +up at the same address with identical inputs, the step runs once and both use the result. If they +arrive with different inputs, that is an error in the workflow (`CacheAddressReuseError`), which almost +always means two scheduled runs were given the same override prefix. + +## Clearing the cache + +`fxtr cache list` shows the cache, and `fxtr cache clear` forgets entries so that the next job to +ask for them runs the steps afresh. Jobs that already used a cleared result keep it. + +```bash +uv run fxtr cache list # everything +uv run fxtr cache list 'qa-v1/answer[model:*,prompt:*]' # a pattern: * matches any key value +uv run fxtr cache clear qa-v1/answer # an address and everything under it +uv run fxtr cache clear --all --step qa.answer # every entry a given step computed +uv run fxtr cache clear --all # the whole cache +``` + +A pattern is an address in which a mapped key's value is `*`. Give every key of the group, in the +address's order, and quote the pattern in the shell. A `*` anywhere else is refused, since it +doesn't mean what it would in a shell glob. `--step NAME` narrows any selection to the +entries one step computed, and `--state` to entries in one state. From Python, the project client +has `list_cache_entries` and `clear_cache_entries`, taking `prefix=`, `addresses=`, or `conflicts_in=`. + +### Resolving a job's conflicts + +When a job stops on cache conflicts, `fxtr jobs status JOB` groups them into patterns and prints +the commands that fix them: + +```text +job 3fa3df92-…: stopped + reason: cache conflicts at 2 addresses, such as 'qa-v1/answer[model:"small",prompt:"math"]': different requests own them; clear them to run these requests there + cache conflicts at 2 addresses: + qa-v1/answer[model:"small",prompt:*] 2 of 3 entries qa.answer + to list them: fxtr cache list --conflicts-in 3fa3df92-… + to clear them: fxtr cache clear --conflicts-in 3fa3df92-…, then fxtr resume 3fa3df92-… + to rerun overwriting: fxtr resume 3fa3df92-… --overwrite-cache-conflicts +``` + +Each pattern is followed by how many of the entries it matches are in conflict, and the step that +owns them: "2 of 3 entries" means the pattern also matches one entry that is not. From +Python, `status.cache_conflicts` lists every conflict as a `CacheConflict`: its `cache_address`, +the `step_name` of the request that owns the entry, and its `kind`, `"completed"` for a result or +`"running"` for an attempt in progress. + +There are three ways forward, and they suit different situations: + + + + Clear exactly the job's conflicts, then resume. You see each conflict before it goes. + + ```bash + uv run fxtr cache list --conflicts-in 3fa3df92-… + uv run fxtr cache clear --conflicts-in 3fa3df92-… + uv run fxtr resume 3fa3df92-… + ``` + + This is the safest approach and is recommended when you want to avoid clearing expensive work. + A downstream step's inputs only change once the steps before it have rerun, so it may conflict + on the resume, and a pipeline of several stages can take a few rounds. + + + + Resume with `--overwrite-cache-conflicts`. Every conflicting step runs and replaces the old + result, as if you had cleared it first. Steps downstream whose inputs change as a result + conflict in turn and are overwritten by the same resume; steps whose inputs come out the same + reuse their results. The flag applies to that one run. `fxtr run` takes it too, and from + Python, `client.submit`, `client.run`, `client.resume`, `run_job`, and `launch_cli` take + `overwrite_cache_conflicts=True`. + + ```bash + uv run fxtr resume 3fa3df92-… --overwrite-cache-conflicts + ``` + + This can be a useful choice while you are iterating on data or settings on purpose and the old + results are no longer wanted. + + + + Launch again under a new `root`. Every address is fresh, every step runs, and every earlier + result stays exactly as it was. Choose this when the old results are still valuable and you may + want to re-run the original workflow with its original inputs. + + + +Neither overwriting nor clearing touches an entry whose step another job is running at that +moment. Wait for that job, or cancel it, and resolve the conflict afterwards. + +### Reusing another job's results + +Clearing the cache does not mean that the data is lost! +Every job maintains an immutable record of the exact step results that were used when it ran. +If you want to restore the cache to a previous state, you can **adopt** the cache entries that +were recorded in any of the jobs in the database using the command `fxtr cache adopt JOB`. This +puts the results a job took back into the cache at their addresses, so later jobs reuse them. + +You can also use this to import parts of the cache from a separate copy of the experiment database +(e.g. from a collaborator's copy). The command `fxtr jobs import FILE --adopt-cache` both makes the +imported job available and adopts all of its cache entries so that they are used by local jobs. + +## Handling changes to step logic + +fxtr tracks the exact version of code used to run any given step, and associates it with the cache +entry. However, the step fingerprint only covers a step's **name** and +**inputs**, not its code. +Editing a step's body and launching under the same root reuses the cached results even if they are +tagged as generated by an older code version. + +This is a deliberate choice, since it is difficult to predict what code changes can affect behavior, +and it is often useful to significantly refactor code without invalidating previous results. We choose +to simply *track* the differing code versions, and allow users to directly control whether or not +their changes should invalidate the cache. + +To make sure that code changes do not lead to invalid results in the cache, we recommend following one of +the following two best practices. + +### Iterating on step logic during development + +While you are first writing a step, you will often make several breaking changes in a row. After each +breaking change, you can clear the cache for that step so the next job does not pick up a stale result: + +```bash +uv run fxtr cache clear --all --step qa.answer +``` + + + To keep trial runs apart from real results, you can also point the project at a scratch database schema + while experimenting: set `schema` in `fxtr.local.toml` to another name, and set it back for + real runs. See [Before you launch](../guides/running-and-viewing-jobs.mdx#before-you-launch). + + +### Adding configuration to stable steps + +Once a step's logic has stabilized and real results depend on it, or once you have shared your +code with collaborators, you should generally avoid changing its behavior in +place. A change in place forces you to remember to clear the cache and to track which version of +the code produced each result, and the system cannot protect you if you forget. Instead: + +* **To change or extend what a step does, add a configuration argument** rather than editing the + logic. Give the new argument a default that preserves the old behavior. Existing results keep + matching, because their inputs are unchanged, and new results computed with the new option have + different fingerprints and so can never be confused with old ones. + (You can also bundle configuration arguments into a [config entity](entities-and-custom-types.mdx), + and handle configuration changes by adding a field to that entity.) +* **If the change is too large for a flag, give the step a new name**, for example by adding a + version suffix: `qa.answer` becomes `qa.answer_v2`. The new name gets fresh cache entries, old + results stay intact and attributable, and jobs that still reference the old step continue to + work. + +The same reasoning applies to anything that influences a result: model IDs, prompts, sampling +settings, dataset contents. Pass them in as inputs, and the fingerprint tracks them for you. A step +that read a prompt template from a module constant would have the same fingerprint even if the +constant changed, and that is exactly the situation the cache cannot detect. diff --git a/plugins/fxtr/skills/fxtr/references/concepts/overview.mdx b/plugins/fxtr/skills/fxtr/references/concepts/overview.mdx new file mode 100644 index 0000000..e217f05 --- /dev/null +++ b/plugins/fxtr/skills/fxtr/references/concepts/overview.mdx @@ -0,0 +1,137 @@ +--- +title: "Core concepts" +description: "How projects, steps, workflows, jobs, arrays, and entities fit together." +--- + +## The mental model + +A fxtr experiment has two halves: + +* The **experiment code** is a program describing the computation you want to run: which steps + to call, on which data, in what order. Experiment code is stored in a git repo (the "project repo"), + and serves as the source of truth of what the experiment does. +* The **job record** is what actually happened when the experiment code ran, including a trace of + all steps and workflows that were executed, along with their inputs and results. Job records are + stored in the experiment database, and can also be exported to files (or imported from files). + +When you launch an experiment, fxtr runs the code and builds the record as it goes, as a graph +connecting every input, operation, and result. The record lives in the project's database, so +you can inspect it in the viewer, resume it after an interruption, and reuse its expensive results +in later runs. + +The concept pages that follow build this picture up one layer at a time: + +1. [Workflows and steps](workflows-and-steps.mdx): how the experiment code is + structured, and the rules each kind of function follows. +2. [Arrays and parallel computations](arrays-and-parallel-computations.mdx): the data + model, and how a step is mapped over data to express a sweep. +3. [Entities and custom types](entities-and-custom-types.mdx): stored records with + identities of their own, for transcripts, verdicts, configurations, and reports. +4. [Managing the step cache](managing-the-step-cache.mdx): how results are reused + across jobs, and how to control that. +5. [Visualization and custom renderers](visualization-and-custom-renderers.mdx): how + the viewer draws a job, and how to teach it about your types. + +## Glossary + + + A uv project whose Python package defines steps and workflows. It's configured by + `[tool.fxtr]` in `pyproject.toml`, and by `fxtr.local.toml`, which says where this machine + reaches the project's Postgres database. + + + + One unit of computation (`@step`), such as a model call, a judgment, or a reduction. A step + receives concrete data, returns a result, and can't schedule further work. Its result is + cached. See [Workflows and steps](workflows-and-steps.mdx). + + + + Deterministic code (`@workflow`) that schedules steps and child workflows, building the + experiment's graph. A workflow works with handles to data rather than the data itself, and can + wait for a result when it needs one to decide what to do next. + + + + A collection of values of one schema, indexed by named dimensions such as `model` or + `prompt`. Every input and output of a step or workflow is an array; a single value is an array + with no dimensions. See [Arrays and parallel computations](arrays-and-parallel-computations.mdx). + + + + A workflow's reference to an array that may not be computed yet, such as a step's future + result. Its dimensions and schema are known immediately, so the workflow can pass it on to + further steps before the data exists. + + + + A filter, projection, aggregation, or join over arrays, written with the query builder or in + SQL. In a step a query runs at once; in a workflow it becomes a node of the graph that the + runner computes when its inputs are ready. + + + + A stored record, such as a conversation or a judge's verdict, with a type and an ID derived + from its content. Arrays hold references to entities, and the viewer links to them. See + [Entities and custom types](entities-and-custom-types.mdx). + + + + One durable execution of a root workflow on a set of input arrays. A job records its progress + as it runs, so it can be inspected while running and resumed if it stops. See + [Running and viewing jobs](../guides/running-and-viewing-jobs.mdx). + + + + A named, typed, editable array kept in the project's database. A launch takes an immutable + snapshot of it to pass as a job input, so a dataset can be curated over time while every job + records exactly the version it ran on. + + + + A human-readable address for each workflow, step, and source array within a job, such as + `root/evaluate/sample[task:"a"]`. It is built from the names you give each operation. + + + + Where a step's result is stored: its pathname, unless you override it. A step reuses a cached + result when its address holds one computed from the same step and inputs. See + [Managing the step cache](managing-the-step-cache.mdx). + + + + A web app, started with `fxtr view`, that shows a project's jobs: each workflow's graph, the + arrays flowing through it, and the entities they reference, through renderers you can + customize. See [Watch jobs in the viewer](../guides/running-and-viewing-jobs.mdx#watch-jobs-in-the-viewer) + and [Visualization and custom renderers](visualization-and-custom-renderers.mdx). + + +## Why steps and workflows are separate + +The split lets fxtr resume jobs and show how each result was produced: + +* Steps hold the expensive, nondeterministic work, such as sampling a model, calling an API, or + drawing a random number. fxtr stores every step's result, so this work isn't repeated when a job + resumes or a later job asks for the same thing. +* Workflows hold the cheap, deterministic decisions about what to run. A workflow makes the same + decisions given the same inputs and results, so fxtr can replay it against the recorded results + to rebuild the graph, and check that the rebuilt graph matches. + +Between the stored results and the recorded graph, fxtr knows which step produced each value, +what that step received, and which workflow asked for it. + +## Hermeticity + +For the job record to be a faithful account of an experiment, the experiment code has to depend +on nothing the record doesn't capture. So steps and workflows should be **hermetic**: they read no files +from the machine they run on, consult no machine-specific state, and reach the outside world only +through stateless services such as model APIs, using credentials supplied by the environment. Code +written this way could run in a sandbox on another machine and produce an equivalent result, which +is exactly what happens when a colleague resumes your job or a later job reuses your cached +results. + +The place for reading local data is the **launcher**, the script that starts a job. It reads and +validates files, turns them into typed arrays, and passes them in as the root workflow's inputs, +or stores them as project arrays for jobs to snapshot. Either way the data becomes part of the +record. [Workflows and steps](workflows-and-steps.mdx#hermeticity-and-handling-external-state) +spells out the rules. diff --git a/plugins/fxtr/skills/fxtr/references/concepts/visualization-and-custom-renderers.mdx b/plugins/fxtr/skills/fxtr/references/concepts/visualization-and-custom-renderers.mdx new file mode 100644 index 0000000..b122652 --- /dev/null +++ b/plugins/fxtr/skills/fxtr/references/concepts/visualization-and-custom-renderers.mdx @@ -0,0 +1,108 @@ +--- +title: "Visualization and custom renderers" +description: "How the viewer draws a job, and how renderers teach it to show your own types, steps, and results." +--- + +fxtr's web **viewer** shows a job as it runs and after it finishes: the graph of workflows and +steps, the arrays flowing between them, and the entities they reference (see +[Watch jobs in the viewer](../guides/running-and-viewing-jobs.mdx#watch-jobs-in-the-viewer)). Everything has a +default rendering, so a new project gets a usable viewer with no extra work. You can customize the visualization by writing **renderers**, +small React components that take over how a particular entity type, step, or workflow is shown, +which fxtr loads into the viewer alongside its own. + +## Renderers, slots, and subjects + +The viewer decides how to draw each thing by looking up a renderer for it. A renderer is +registered for a **slot**, which is a place in the viewer where something is drawn, and a +**subject**, which says what kind of thing it draws there: + +| Slot | Where it appears | Subject | +| -------------------- | ----------------------------------------------------------------- | ---------------------------------------- | +| `EntityPanel` | An entity opened in the detail pane | The entity's type tag | +| `EntityInlineLink` | An entity where it is referenced, as a link | The entity's type tag | +| `EntityHoverPreview` | A compact card shown when a link is hovered | The entity's type tag | +| `InvocationPanel` | A full-pane view of one or more invocations of a step or workflow | The step's or workflow's registered name | + +For entities, the subject is the `fxtr_type` string from the entity's definition, which is why +every entity type needs one. For invocations, it is the `name=` given to `@step` or `@workflow` (or, by default, the Python function's module-qualified name). +If no renderer is registered for a subject, the viewer falls back to generic ones that can be used with any value. +Several renderers may target the same slot and subject, in which case the viewer offers a switcher. + +A renderer is worth writing when an entity holds a lot of data, or when some of its fields deserve +more emphasis than others. A renderer for an entity type is an ordinary component that receives +the entity's ID: + +```tsx views/src/answer.tsx +import { EntityLink, unwrap, useEntity, type EntityID, type EntityViewProps } from "fxtr-view"; +import { FieldList, Kind, layout } from "fxtr-view-kit"; + +interface AnswerData { + config: EntityID; + question: string; + text: string; +} + +export function AnswerPanel({ id }: EntityViewProps) { + const answer = unwrap(useEntity(id as EntityID)); + return ( +
+
Answer
+

{answer.question}

+

{answer.text}

+ ]]} /> +
+ ); +} +``` + +`unwrap(useEntity(id))` returns the entity's decoded data, or suspends until it arrives, so the +component body can assume the data is there and leave the loading state to the viewer. The data is +the entity's stored form: a map with the entity's fields, where references to other entities arrive +as `EntityID`s that `EntityLink` turns into links. + +The project's **bundle** collects its renderers, registering each with its slot, its subject, an +ID, and the component: + +```ts views/src/bundle.ts +import { EntityPanel, InvocationPanel, renderer, type ViewBundle } from "fxtr-view"; +import { singleInvocation } from "fxtr-view-kit"; +import { AnswerPanel } from "./answer"; +import { Overview } from "./overview"; + +const bundle: ViewBundle = { + renderers: [ + renderer(EntityPanel, "org.example.qa.Answer.v1", "my-project.AnswerPanel", AnswerPanel), + renderer(InvocationPanel, "qa.evaluate", "my-project.Overview", singleInvocation(Overview)), + ], +}; + +export default bundle; +``` + +Renderers are built with two libraries. **`fxtr-view`** is the small, stable contract between a +renderer and the viewer: the slots, `renderer(...)`, and hooks for loading data and linking to +entities. **`fxtr-view-kit`** has building blocks that match the viewer: layout classes, chart +parts and controls, number formats, and adapters for showing invocations and loading many entities +at once. The [renderer reference](../reference/renderers.mdx) describes both, with the rules for +registering renderers. + +## Results overviews + +The most valuable renderer in most projects is an `InvocationPanel` for the **root workflow**, like +the `Overview` registered above. The viewer offers it as a full-pane view of the whole run, in +place of the trace, and it is where you present the experiment's results: the charts, tables, and +comparisons someone would want to see first. It usually reads a **report** entity that the root +workflow returns, which references every result array worth presenting, so one load gives the +overview everything it needs (see [Read results](../guides/running-and-viewing-jobs.mdx#read-results)). +`fxtr new --example` writes a small one, `views/src/overview.tsx`, and +[Designing views](../guides/designing-views.mdx) covers what to show in one and how to draw it. + +## Renderer packages and view bundles + +A project's renderers live in a TypeScript package of their own, `views/` in a project created by +`fxtr new`, which is built into a **view bundle** that the viewer loads at runtime. The viewer +hands the bundle its own copies of React and `fxtr-view`, so renderers share its state and context, +and nothing about a project is compiled into the viewer, so one viewer serves every project. While +it runs, `fxtr view` rebuilds the bundle whenever the renderers' sources change, and open pages +switch to each new build in place. [The renderer package](../reference/renderers.mdx#the-renderer-package) +covers its parts, building it, and what to do when the bundle doesn't load. diff --git a/plugins/fxtr/skills/fxtr/references/concepts/workflows-and-steps.mdx b/plugins/fxtr/skills/fxtr/references/concepts/workflows-and-steps.mdx new file mode 100644 index 0000000..aa6a717 --- /dev/null +++ b/plugins/fxtr/skills/fxtr/references/concepts/workflows-and-steps.mdx @@ -0,0 +1,359 @@ +--- +title: "Workflows and steps" +description: "How an experiment is written: workflows describe the structure, steps do the work." +--- + +In a fxtr project you launch **jobs**, and each job executes one **workflow**. A workflow +describes the structure of an experiment: which computations to run, on which data, in what +order. It does not do the expensive work itself. Instead it schedules **steps**, and a step is +where the expensive or nondeterministic computation happens: sampling a model, judging an answer, +drawing a random number, crunching a large array. + +A job is therefore a tree: one root workflow, which schedules steps and child workflows, which in +turn schedule steps. All the work happens at the leaves. + +```mermaid +--- +config: + markdownAutoWrap: false + flowchart: + wrappingWidth: 600 +--- +flowchart LR + J["job"] --> W["root workflow
root"] + W --> S1@{ shape: procs, label: "step
root/answer[model:M,prompt:P]" } + W --> S2@{ shape: procs, label: "step
root/score[model:M,prompt:P]" } + W --> C@{ shape: procs, label: "child workflow
root/summarize[model:M]" } + C --> S3@{ shape: procs, label: "step
root/summarize[model:M]/mean" } + classDef job fill:#e7e5e4,stroke:#78716c,color:#1c1917 + classDef workflow fill:#dbeafe,stroke:#2563eb,color:#1e3a8a + classDef step fill:#dcfce7,stroke:#16a34a,color:#14532d + class J job + class W,C workflow + class S1,S2,S3 step +``` + +Workflows are required to be functionally pure and deterministic: given the same inputs and the +same step results, a workflow must make the same decisions. That property is what lets fxtr +rebuild a job's graph by replaying the workflow, resume a job after an interruption, and check that +nothing changed under it. + +Steps, on the other hand, are allowed to be nondeterministic, and may involve expensive work such +as querying a language model. Fxtr executes each step once and saves its result in the **step cache**, +so that previously computed results can be re-used. + +## Workflows + +A workflow is an async Python function decorated with `@workflow`. It receives a context and a set +of named inputs, and it returns a result. Its inputs and result are **arrays**, collections of +values indexed by named dimensions, which +[Arrays and parallel computations](arrays-and-parallel-computations.mdx) covers in +depth. Inside the workflow you mostly do not touch the arrays' data. You work with +**handles** to them, and you pass those handles between the operations you schedule. + +Here is a workflow that scores a set of prompts with a model, then asks a child workflow to +summarize the scores: + +```python +from typing import Annotated + +from fxtr.experiment.handles import ArrayHandle +from fxtr.experiment.workflows import WorkflowContext, workflow + + +@workflow(name="qa.evaluate") +async def evaluate( + context: WorkflowContext, + prompts: Annotated[ArrayHandle[str], "[prompt: str]"], + models: Annotated[ArrayHandle[BoundID[ModelConfig]], "[model: str]"], +) -> Annotated[ArrayHandle[float], "[model: str]"]: + """Answer every prompt with every model and summarize each model's scores.""" + answers = context.run_step( + "answer", answer, {"config": models, "question": prompts}, map_over=["model", "prompt"] + ) + scores = context.run_step("score", score, {"answer": answers}, map_over=["model", "prompt"]) + return context.run_child_workflow("summarize", summarize, {"scores": scores}, map_over=["model"]) +``` + +Three things are happening: + +* `context.run_step(name, step, inputs, ...)` schedules a step and immediately returns a handle to + its future result. It does not wait. The `map_over` argument says which dimensions to run the + step once per key of, so `answer` is called once for every model and prompt pair. +* Passing the `answers` handle to the `score` step is what makes `score` depend on `answer`. fxtr + draws an edge between them in the viewer, and will not start `score` until the answers exist. +* `context.run_child_workflow(...)` works the same way for a workflow. Any workflow can be the root + of a job, and any workflow can be a child of another, so you can compose an experiment out of + reusable parts. + +When the body returns, the workflow's job is done: it has described a graph. The runner then +works through the graph, running each step whose inputs are ready, until every result exists. + +### Determinism + +fxtr runs a workflow body more than once. It replays the body to resume a job after a crash, and +by default it also runs the body a second time at the end of every attempt to verify that the body +made the same decisions. A workflow must therefore be cheap and deterministic: + +* Do anything expensive, random, or dependent on the outside world in a step. +* Never generate random numbers, UUIDs, or timestamps in a workflow. +* Give every scheduled operation a stable name. Names are how fxtr matches a replayed operation + with the recorded one. +* Fix datasets and settings before the job starts and pass them in as inputs. + +If a replayed body builds a different graph under the same names, the attempt fails with a +`DeterminismError` rather than silently continuing with a mismatched record. + + + Workflow determinism is only required within a single job. If you change your workflow code, you can simply start a new job to execute it from scratch. The step cache is shared across jobs, so this still allows you to re-use expensive work. + + +### Handles + +A handle (`ArrayHandle`) is a reference to an array that may not have been computed yet. Its +**type**, the dimensions and the kind of value it holds, is known the moment you receive it, because +fxtr derives the type of every operation from the types of its inputs. That is why the workflow +above can pass `answers` to the next step before any model has been called: the type checks +happen while the graph is being built, and a mismatch is an error at scheduling time rather than +after an hour of sampling. + +You can also ask a handle for its data, with `await handle.observe()` for the whole array or +`await handle.item()` for a single value. This pauses the workflow until the result exists, and it +is how a workflow makes a data-dependent decision (see +[Data-dependent control flow](#data-dependent-control-flow)). +Mapping a step over a handle's dimensions, and the other ways of combining handles, are covered in +[Arrays and parallel computations](arrays-and-parallel-computations.mdx#mapping-logic-over-arrays). + +## Steps + +A step is where the real work happens. It is an async function decorated with `@step` that +receives concrete data, computes something, and returns a result. Model calls, LLM judges, +simulations, external API queries, random sampling, and heavy numerical work all belong in steps. +So does any deterministic transformation whose result you want cached or whose place in the graph +you want to see in the viewer. + +Here is a step that asks a language model one question, using the +[behaviors](../../../behaviors/references/index.mdx) library: + +```python +import anyio +from behaviors.implementations.models import ( + DEFAULT_FAILURE_CATEGORIES_TO_RETRY, + generate_with_retries, +) +from behaviors.implementations.models.openai_responses import OpenAIResponsesAPI +from behaviors.types import ContentText, ModelCallFailure, ModelRequest, RetryConfig, UserMessage + +from fxtr.experiment.steps import StepContext, step + + +@step(name="qa.answer") +async def answer(context: StepContext, config: ModelConfig, question: str) -> Answer: + """Ask the model one question and keep its reply.""" + request = ModelRequest( + context=(UserMessage(content=(ContentText(text=question),)),), + tools=(), + tool_choice=None, + parallel_tool_calls=None, + sampling=config.sampling, + ) + api = OpenAIResponsesAPI(config.model_name) + try: + response = await generate_with_retries( + api, request, retry=RetryConfig(), + failure_categories_to_retry=DEFAULT_FAILURE_CATEGORIES_TO_RETRY, + ) + finally: + with anyio.CancelScope(shield=True): + await api.aclose() + if isinstance(response, ModelCallFailure): + raise RuntimeError(f"model call failed ({response.category}): {response.description}") + return Answer(question=question, text=answer_text(response.output)) +``` + +`ModelConfig` and `Answer` are entity types the project defines (see +[Entities and custom types](entities-and-custom-types.mdx)), and `answer_text` is a +small helper that joins the text blocks of the reply. The details of building requests, holding +multi-turn conversations, and judging outputs are in +[Calling language models](../../../behaviors/references/guides/language-models.mdx). + +### How steps run and get cached + +To run a step, schedule it from a workflow with `context.run_step`. Scheduling gives the +invocation an **address** in the job and records what inputs it received, and it is what connects +the step to the cache, to the viewer, and to the job's record. (Don't call the step function +directly from your own code, this will not be tracked by fxtr!) + +Before running a step, fxtr looks at that address in the project's **step cache**. If a result +computed by the same step from the same inputs is already there, the job simply +reuses the result. If there is no result there, the step runs and its result is stored at the address. This is what +lets you edit a workflow, launch it again, and pay only for the new work. When the cache holds a result computed +from *different* inputs, the job either pauses for review or overwrites the cache, depending on the configuration. +[Managing the step cache](managing-the-step-cache.mdx) explains this in more detail. + +### Failures: return or raise? + +A step that raises an exception is treated as having hit a **transient** problem. Its attempt +fails, nothing is cached, work that does not depend on it keeps going, and the job ends without +succeeding. When you resume the job, the step runs again. This is the right behavior for a network +error, a rate limit that outlasted the retries, or a crashed subprocess. + +A step that **returns** a value is treated as having succeeded, and its result is cached like any +other, even if that result records an error. So the rule is: + +* **Return** for an outcome every attempt would reproduce: an input with nothing to grade, a reply + that did not contain the expected answer, a conversation that hit the context limit. Represent the + outcome in your result type, for example with a nullable score or an error field. +* **Raise** for a failure a later attempt might not hit. + + + Caching does not make side effects happen exactly once. An attempt that is interrupted after it + sent a request but before its result was saved will run again when the job resumes. + + +### Steps should be pure, even if nondeterministic + +A step may be nondeterministic. Two samples from the same model with the same prompt differ, and +that is fine, because fxtr records whichever one the step returned and never recomputes it. But a +step should still be **functionally pure**: its result should depend only on its inputs, plus the +external services it calls through stateless APIs. It should not read files from the machine it +happens to run on, consult environment-specific state, or write anything anywhere except through +its return value and the entities it stores. The reasons are spelled out in +[Hermeticity](#hermeticity-and-handling-external-state) below. + +## Contexts and signatures + +### Signatures + +Every step and workflow has a **signature**: the array type of each input and of the result. fxtr +reads it from the function's type annotations. A few conventions cover most cases: + +| Annotation | Means | +| ---------------------------------------------- | -------------------------------------------------------------------------- | +| `question: str`, `-> float` | A single value. The step body receives and returns the bare value. | +| `Annotated[Array[float], "[sample: int]"]` | A concrete array with a `sample` dimension. Used in steps. | +| `Annotated[ArrayHandle[str], "[prompt: str]"]` | A handle to an array with a `prompt` dimension. Used in workflows. | +| `ScalarHandle[int]` | A handle to a single value. Short for `Annotated[ArrayHandle[int], "[]"]`. | + +The string inside `Annotated` names the dimensions, which no Python type can express on its own. +`Array[T]` or `ArrayHandle[T]` alone says what the values are but not what the dimensions are, so +always wrap one in `Annotated[..., "[...]"]`, including `"[]"` for a scalar. + +For an input that is a reference to an [entity](entities-and-custom-types.mdx), a +stored record with an ID of its own, the annotation also decides what the body receives: + +* `config: ModelConfig` loads the entity before the body runs. This is the common case. +* `config: BoundID[ModelConfig]` passes a typed reference without loading it, which is what you + want when the step's result should point back at the input. +* `config: EntityID` passes the bare ID. +* `Annotated[Array[BoundID[Document]], "[doc: str]"]` passes a whole array of references. + +A step declared to return an entity type, like `-> Answer` above, can return the entity itself. +fxtr stores it and records its reference as the step's result. A step declared to return a struct +returns its `TypedDict`, and a step whose result has dimensions returns an `Array`. + +A workflow's parameters are usually handles. Annotating one with another type makes the runner +wait for that input and pass its data instead: `Array[...]` passes the array, a single-value type +such as `int` passes the value, and an entity type passes the loaded entity. A workflow returns a +handle, or, for a single-valued result, a plain value or entity, which fxtr records as a source +array named `$return`. Returning the handle of the step that computed a result keeps the two +linked in the viewer. + +When types are easier to state than to annotate, give them to the decorator instead: +`@step(inputs={"prompt": str}, output=int)`, writing a type with dimensions as a pair such as +`("[sample: int]", float)`. Annotations given as well must agree with these. For the rare step or +workflow whose result type depends on its inputs' types, a custom `Signature` (`signature=`) +computes it. + +### The context argument + +The first parameter of every step and workflow is its **context**, and it must be named +`context`: the decorators' types require that name, so a type checker such as Pyright rejects any +other. A `StepContext` and a `WorkflowContext` both let you: + +* **Store and load entities.** `ref = await context.store(entity)` returns a reference, and + `await context.load(ref)` turns one back into the object. Entities stored this way become part + of the job's record. +* **Log.** `context.log("scored 40 of 100")` writes a line to the attempt's log, which the viewer + shows alongside the invocation. + +A `WorkflowContext` additionally has the scheduling operations: `run_step`, +`run_child_workflow`, `add_source_array` for introducing fixed data into the graph, and `combine` +and `stack` for putting handles side by side. A `StepContext` exposes the step's `cache_address`, the +address its result is stored at. + +### Registered names + +Each step and workflow is registered under a name: the `name=` given to `@step` or `@workflow`, +or by default the function's module and qualified name, such as +`my_project.experiment.count_words`. The name identifies the function in a job's record and to +the viewer's renderers, a resumed job finds the function by it, and the step cache matches a step +by its name and inputs (see [Managing the step cache](managing-the-step-cache.mdx)). +Setting `name=` keeps the name stable when the function moves to another module. A function +defined in a script that is run directly (in `__main__`) has no stable default, so it must set +one. Importing a module is what registers its functions, which is why a project lists its +experiment modules in `[tool.fxtr] modules` (see +[Project configuration](../reference/project-configuration.mdx)). + +## Data-dependent control flow + +Because a workflow is ordinary Python, it can look at a result and decide what to schedule next. +Observing a handle waits for its value, so the workflow pauses at that point, then continues +building the graph with the value in hand. The classic shape is a refinement loop that continues +until a judge is satisfied: + +```python +current = draft +for round in range(max_rounds): + review = context.run_step(f"review_{round}", review_draft, {"draft": current}) + proposal = await review.field("proposal").collect().item() + if proposal is None: + break + current = context.add_source_array(f"draft_{round + 1}", Array.scalar(proposal, BoundID[Draft])) +return current +``` + +Each round's step gets a stable name that includes the round number, so replaying the workflow +rebuilds the same sequence. The stored job records how many rounds each input actually took, which +a static graph could not express. + + + Observe a result only to decide what to schedule next. If you observe an array, transform it in + Python, and add the transformed array back as a source array, the viewer loses the edge between + the two, and the transformation is not cached. Put the transformation in a step or a + [query](arrays-and-parallel-computations.mdx#querying-filtering-and-aggregating-arrays) + that takes the original handle instead. + + +## Hermeticity and handling external state + +Everything inside a workflow function or a step function should be **hermetic**: it should be +possible to run it in a sandbox outside your environment, on a different machine, and get an +equivalent result. The reason is that fxtr's record of a job is meant to be a complete account of +what was computed. It stores the inputs every step received, the code commit the job ran at, and +every result. If a step quietly read a file from your laptop, that file is not in the record, the +step's cache entry cannot be trusted (the address and fingerprint say nothing about the file's +contents), and the job cannot be resumed by anyone who lacks the file. + +Concretely, inside steps and workflows: + +* **Don't read local files.** No `open(...)`, no paths relative to the project, no datasets loaded + from disk. The launcher is the only place that reads files. +* **Don't depend on the machine.** No environment variables that change behavior, no hostnames, no + reliance on what else is installed. Credentials for model providers are the one accepted + exception: they live in the environment, never in inputs or entities, and whoever runs the job + must supply them. +* **Call only stateless services.** Model APIs and similar request-response services are fine. A + service whose answer depends on state your step changed earlier is not, because a replay would see + different state. +* **Don't write anywhere except the record.** A step's output is its return value and the entities + it stores. Writing files or updating external databases from a step leaves things behind that the + record does not know about. + +The recommended pattern for local data is to read it in the launcher, convert it to a typed +array, and pass it in as an input of the root workflow. For a dataset you curate over time, store +it in the database as a **project array** and launch jobs against snapshots of it. Both are +described under [Creating arrays](arrays-and-parallel-computations.mdx#creating-arrays). +Either way, the data ends up in the job's record, every step that uses it has a fingerprint that +reflects its contents, and anyone with access to the project's database can resume or reproduce +the job. diff --git a/plugins/fxtr/skills/fxtr/references/first-experiment.mdx b/plugins/fxtr/skills/fxtr/references/first-experiment.mdx new file mode 100644 index 0000000..315f769 --- /dev/null +++ b/plugins/fxtr/skills/fxtr/references/first-experiment.mdx @@ -0,0 +1,222 @@ +--- +title: "Your first experiment" +description: "Write a small experiment with a mapped step and a reduction, launch it, and see how caching works." +--- + +This tutorial starts with an empty project and builds an experiment that counts the words in +each prompt in a dataset, then averages the counts. + +This assumes you have completed [Installation](installation.mdx) and have a database running. + + + **Claude can write this experiment for you.** With fxtr's + [agent skills](installation.mdx#install-the-agent-skills) installed, you can describe it in + plain language, such as "in a new fxtr project, count the words in each of these + prompts, average the counts, and launch it", and review the result in the viewer. + + +## Create a project and connect it to the database + +Create a new project in a new or empty directory, and connect it to your database. The command +below uses the local database from [Installation](installation.mdx#start-postgres); if you +use another Postgres server, put its URL instead. + +```bash +uvx fxtr new ~/code/my-project +cd ~/code/my-project +uv sync +(cd views && pnpm install) +uv run fxtr init --database-url postgresql+asyncpg://fxtr:fxtr@localhost:55433/fxtr +``` + +## Write the experiment + +Create `src/my_project/word_count.py`: + +```python src/my_project/word_count.py +from typing import Annotated + +from fxtr.entity_defns.array import Array +from fxtr.experiment.handles import ArrayHandle +from fxtr.experiment.steps import StepContext, step +from fxtr.experiment.workflows import WorkflowContext, workflow + + +@step(name="word_count.count") +async def count_words(context: StepContext, prompt: str) -> int: + """Count the whitespace-separated words in a prompt.""" + return len(prompt.split()) + + +@step(name="word_count.mean") +async def mean_length( + context: StepContext, + counts: Annotated[Array[int], "[prompt: str]"], +) -> float: + """Average the word counts across prompts.""" + values = list(counts.values()) + return sum(values) / len(values) + + +@workflow(name="word_count.experiment") +async def experiment( + context: WorkflowContext, + prompts: Annotated[ArrayHandle[str], "[prompt: str]"], +) -> Annotated[ArrayHandle[float], "[]"]: + """Measure the length of each prompt, then average the lengths.""" + counts = context.run_step("count_words", count_words, {"prompt": prompts}, map_over=["prompt"]) + return context.run_step("mean_length", mean_length, {"counts": counts}) +``` + +In this code: + +* `count_words` takes one prompt and returns one number. `mean_length` takes a whole array of + counts, indexed by the `prompt` dimension, and returns one number. +* `run_step` schedules a step and returns a *handle* to its future result without waiting for + it. Passing `counts` to the second step makes it depend on the first. +* `map_over=["prompt"]` calls `count_words` once per prompt and stacks the results into an array + with the same `prompt` dimension. The second `run_step` has no `map_over`, so `mean_length` + runs once on the whole array. +* A plain `str` or `int` annotation means a single value. `Annotated[..., "[prompt: str]"]` also + names the array's dimensions, which a Python type can't express. +* Each docstring appears on its step's or workflow's card in the viewer. + +Then register the module in `pyproject.toml`. + +```toml pyproject.toml +[tool.fxtr] +modules = ["my_project.word_count"] +``` + +## Launch it + +Create a launcher at the project root. It builds the input dataset and runs the workflow as a +job: + +```python launch_word_count.py +import anyio + +from fxtr.entity_defns.array import Array +from fxtr.project.running import open_local_client, run_job +from my_project.word_count import experiment + + +async def main() -> None: + prompts = Array.from_items( + [ + (("greeting",), "Hello there"), + (("question",), "Why is the sky blue?"), + (("request",), "Please summarize this article in three sentences."), + ], + ("[prompt: str]", str), + ) + async with open_local_client(__file__) as client: + result = await run_job(client, experiment, {"prompts": prompts}, root="word-count-v1") + print(f"mean length: {result.item():.2f} words") + + +if __name__ == "__main__": + anyio.run(main) +``` + +`Array.from_items` takes `(key, value)` pairs and a type: here, one `prompt` dimension with string +keys, holding string values. `root` names the job's root workflow, and, as you'll see below, +decides which earlier results the job can reuse. + +Commit, then launch, with `uv run fxtr view` running in another terminal. + +```bash +git add . && git commit -m "Add the word count experiment" +uv run python launch_word_count.py +``` + +```text +job 3eb75a59-911e-44af-a148-411ac5dcd8d1: running +watch it at http://127.0.0.1:8000/#job=3eb75a59-911e-44af-a148-411ac5dcd8d1 +job 3eb75a59-911e-44af-a148-411ac5dcd8d1: succeeded +... +mean length: 4.67 words +``` + +The launch opens the job in a new tab of the viewer: the `prompts` input feeding three +`count_words` calls, one per prompt, feeding `mean_length`. + +## Rerun it + +Launch it again, unchanged. The job succeeds straight away, because every step finds its result +in the cache. You can see the entries with `fxtr cache list`: + +```text +completed word-count-v1/count_words[prompt:"greeting"] (word_count.count) +completed word-count-v1/count_words[prompt:"question"] (word_count.count) +completed word-count-v1/count_words[prompt:"request"] (word_count.count) +completed word-count-v1/mean_length (word_count.mean) +``` + +Each entry is at a **cache address** built from the job's `root`, the step's name in the +workflow, and, for a mapped step, the key it was called with. A step reuses a cached result when +its address holds a result computed from the same step and the same inputs. + +## Change an input + +Now edit the `question` prompt, commit, and launch again under the same `root`. The job stops: + +```text +job 3fa3df92-2edb-4088-9e1f-a99d28b7ec77: stopped + reason: cache conflict at 'word-count-v1/count_words[prompt:"question"]': a different request owns the address; clear it to run this request there + cache conflicts at 1 address: + word-count-v1/count_words[prompt:"question"] 1 address word_count.count + to list them: fxtr cache list --conflicts-in 3fa3df92-2edb-4088-9e1f-a99d28b7ec77 + to clear them: fxtr cache clear --conflicts-in 3fa3df92-2edb-4088-9e1f-a99d28b7ec77, then fxtr resume 3fa3df92-2edb-4088-9e1f-a99d28b7ec77 + to rerun overwriting: fxtr resume 3fa3df92-2edb-4088-9e1f-a99d28b7ec77 --overwrite-cache-conflicts +Traceback (most recent call last): + ... +fxtr.project.client.JobNotFinishedError: Job 3fa3df92-2edb-4088-9e1f-a99d28b7ec77 is stopped +``` + +The address `word-count-v1/count_words[prompt:"question"]` already holds a result computed from +the old prompt. fxtr won't reuse that result for the new prompt, and doesn't replace a completed +result unless you ask, so this is a **cache conflict**: the step doesn't run, and the job stops +once nothing else can. The unchanged prompts still reuse their results. `run_job` raises +`JobNotFinishedError` for any job that doesn't succeed, so your launcher can tell. + +The simplest way forward is to clear exactly the conflicting entries and resume the job: + +```bash +uv run fxtr cache clear --conflicts-in 3fa3df92-2edb-4088-9e1f-a99d28b7ec77 +uv run fxtr resume 3fa3df92-2edb-4088-9e1f-a99d28b7ec77 +``` + +`count_words` runs again for the new prompt. If its new count changes the mean, the resume stops +again, this time at `mean_length`, whose input only changed once `count_words` had rerun. Clear +and resume once more and the job succeeds. Clearing only affects later requests. + +When you are changing data on purpose and don't need to look at each conflict first, +`uv run fxtr resume JOB --overwrite-cache-conflicts` does the same in one go, rerunning every +conflicting step and replacing its old result, downstream steps included. A third option is to +launch under a new `root`, which leaves every earlier result intact. +[Resolving a job's conflicts](concepts/managing-the-step-cache.mdx#resolving-a-jobs-conflicts) +compares the three. + + + Changing a step's **code** does not invalidate its cached results: the cache matches a step by + its name and inputs, not its implementation. After changing what a step computes, launch under + a new `root` or clear its entries. See + [Handling changes to step logic](concepts/managing-the-step-cache.mdx#handling-changes-to-step-logic). + + +## Next steps + + + + The rules workflows and steps follow, and why they are separate. + + + + Sweep over models, prompts, and samples, and reduce over any dimension. + + + + Use behaviors to hold conversations and judge outputs in your steps. + + diff --git a/plugins/fxtr/skills/fxtr/references/guides/designing-views.mdx b/plugins/fxtr/skills/fxtr/references/guides/designing-views.mdx new file mode 100644 index 0000000..081626b --- /dev/null +++ b/plugins/fxtr/skills/fxtr/references/guides/designing-views.mdx @@ -0,0 +1,319 @@ +--- +title: "Designing views" +description: "How a renderer's views should look and behave in the viewer: surfaces, type, color, charts, links, controls, and how to check them." +--- + +A project's renderers draw inside the experiment viewer, beside the viewer's own cards, tables, +and panes. A view has two parts. Its frame (headings, text, tables, controls, links) is the +viewer's: the same dense type, the same controls, and links and selection that behave like the +viewer's own, so the view does not look out of place. Its charts are figures set in that frame, +with a visual language of their own: their colors mean what their legends and labels say. This +guide is how the viewer looks and behaves, and how to build charts and controls that sit well in +it with `fxtr-view` and `fxtr-view-kit`. The [renderer reference](../reference/renderers.mdx) +covers registering renderers, loading data, and the kit's API. + +A project written with `fxtr new --example` has a complete small example in +`views/src/overview.tsx`: a results overview with a chart that follows everything below. + +## Where a view appears + +| Surface | Renderer | Room | Base type | Surface under it | +| ---------------- | -------------------------------------------- | ------------------------------------------------------------------------------------------------------- | --------- | ---------------- | +| Results overview | `InvocationPanel` of the root workflow | the page, up to 1400px wide, 12px side margins | 13px | `--bg-alt` | +| A pane | `EntityPanel`, or a step's `InvocationPanel` | resizable; can be 300px or less | 12px | `--bg-alt` | +| A card's preview | the node's default view | beside each card while no pane is open, about half the page wide, as tall as the card, scrolling inside | 12px | `--bg-alt` | +| Hover card | `EntityHoverPreview` | 360px wide, at most 300px tall, the full cid below it | 11px | `--bg` | +| Inline link | `EntityInlineLink` | one line inside a pill, cut off with an ellipsis | 11px | the pill | + +A project's own renderer becomes the default view of its node, so the same component may draw at +full page width, in a narrow pane, and in a card's preview. Design for every width from about 320px +to 1400px, and put the finding that matters most at the top left, where every surface shows it. + +The viewer draws the frame. Leave out what it already gives: + +* No padding, background, border, shadow, or font size on the view's root. The view inherits its + surface and its base size, which differ by surface. +* The pane's head already names the entity or node. An `EntityPanel` starts with its kind + (`Transcript`); an invocation view starts with `InvocationHead`. +* The top right corner of an entity's panel holds the viewer's view switcher and info button, + faint until hovered. Keep controls out of that corner. + +## The look + +### Type + +The viewer is dense: most text is 12px or less, and nothing is larger than 14px. Use the type +tokens, never a pixel size of your own. + +| Token | Size | For | +| ------------- | ---- | -------------------------------------------------------------------- | +| `--text-2xs` | 9px | rarely; tiny tags | +| `--text-xs` | 10px | axis ticks, legends, section headings, notes beside a number | +| `--text-sm` | 11px | table headers, captions, tooltips, controls, chips | +| `--text-md` | 12px | body text in panes, table cells | +| `--text-base` | 13px | body text of a results overview (inherited) | +| `--text-lg` | 14px | the one title of a panel (`layout.title`), headline numbers (`Stat`) | + +* Fonts are `--font-sans` (inherited) and `--font-mono`. Monospace is for identifiers: dimension + keys, IDs, argument names, kinds (`Kind` is monospace already). Everything else is sans. No + other family, and no serif or display face. +* A headline number is a `Stat` (14px, weight 600), not a large numeral. +* Section headings are `

`: small, spaced, dim, and the one + uppercase style in the viewer, set by its CSS. Write the heading's text in lower case. Never + write uppercase text in the source. +* Labels, headers, and options are lower case ("model", "mean score"), not Title Case. Prose is + sentence case. +* Numbers use tabular figures and sit right-aligned in columns (`layout.num`). Write them with the + kit's formatters: a true minus sign, grouped thousands, a fixed number of decimals per column, + and `—` for a missing value. + +### Color + +Every color is a design token, written `var(--token)` in CSS, in inline styles, and in SVG fills +and strokes alike. Never write a literal color (`#fff`, `white`, `rgb(...)`, `hsl(...)`): the +viewer has a light and a dark theme, the tokens switch with it, and a literal stays the same in +both. Never set `color-scheme`, never redefine a token, and never paint a background of your own +behind the view. + +For the frame, use the tokens the way the viewer does: + +| Tokens | What the viewer uses them for | +| ---------------------------------------------------------------- | ------------------------------------------------------------------------------------------------- | +| `--fg`, `--fg-dim`, `--fg-faint` | ink: text, then secondary text, then the faintest | +| `--bg`, `--bg-alt`, `--muted-bg` | surfaces: a box's fill, the page, a neutral track or empty cell | +| `--line-soft`, `--line`, `--line-strong` | hairlines, from faintest: row separators; rules, borders, grid lines; axis lines, control borders | +| `--hover`, `--hover-strong` | translucent washes over a hovered row or button, and a pressed one | +| `--select-accent`, `--select-bg`, `--select-fg`, `--select-line` | blue: what a pane shows, and the selection | +| `--picked-bg`, `--picked-fg` | gray: a row picked in a table | +| `--danger`, `--danger-bg` | red: failure | +| `--mapped-*`, `--replica-*` | warm and green: the kinds of a run's dimensions | +| `--role-*-bg` | chat turns, by speaker (the kit's `MessageList` uses them) | + +In a chart, any of these can carry data too: blue for one condition and red for another is fine +when the legend says so. The chart colors are a default palette for data, checked to stay apart +for color-blind readers and readable in both themes: + +* **Categorical** (which series): `--chart-1` to `--chart-8` (blue, orange, aqua, ochre, + magenta, green, violet, red), through `seriesColor(i)`, assigned in order. One series is + `--chart-1`. Never cycle: past eight series, fold the rest into one "other" series + (`CHART_OTHER`, a gray), or split the chart into small multiples. A series keeps its color when + a filter hides others. In a scatter, a map, or small multiples, where any two colors can sit + side by side, the first three stay apart for every reader; past three, label the series + directly. +* **Sequential** (how much): `sequentialFill(t)` for `t` from 0 (near nothing) to 1 (the most), + one blue from pale to deep. Text on it: `sequentialInk(t)`. +* **Diverging** (which side of a baseline): `divergingFill(t)` for `t` from -1 (red) to 1 + (blue), through a gray midpoint at 0. Text on it: `divergingInk(t)`. +* Color does a job another channel is not already doing: bars of one series are all one color, + not colored by their length. +* Identity never rests on color alone: label series directly, or add a `Legend`. + +Where marks need separating, leave real space between them, not a stroke or border painted +`var(--bg)`: views sit on `--bg-alt`, so a `--bg` halo shows as a pale seam. + +### Lines, boxes, and motion + +* Group with space and a `layout.subtitle` heading, not with cards. Never put a card inside a card. + Where a box is needed (a pick's detail), use the kit's `DetailBox`: a 1px `--line` border, 4px + corners, no shadow. +* Corners: a pill's are fully round, a control's or a box's are 4px, a bar's or a cell's 2px. +* Shadows only on what floats (the tooltip). Rules and borders are 1px. +* No animation or transition of your own; the viewer animates only what is running. + +## Charts + +### Choose the form + +First decide what someone most needs to see from the experiment. It may not be the final step's +output; it may come from several steps. Then pick the chart that shows it. A good graph is worth a +thousand words, and a standard bar chart is often not the best one. Make the graphs as clear and +as information-dense as they can be, and get the details right. + +* **Show distributions, not just averages**, when there is room: a dot strip, a histogram, or a + box beside the mean. Density trades against clarity: don't make a graph too busy to read. Add + interactivity to show more detail, not for its own sake. +* **Consistent axes.** Charts of the same kind side by side (two accuracy charts on different + slices) share their ranges and their category order and spacing. Mismatched axes on charts that + look alike mislead. +* **One value axis per chart.** Two measures of different scales are two charts. +* **Two or more categorical variables:** + + * a *table* of X × Y, the baseline for two; + * *small multiples*: one table or chart per value of a third variable Z, for a small Z (about + six or fewer), when the interaction with Z should be visible at a glance; + * *a chart in each cell*: an X × Y table whose cells hold a small chart over Z (a bar, a dot + strip, a sparkline), for a scalar output and a larger Z; + * *coordinate strips*: one strip per variable, each showing the output along that variable + with the others fixed at a current point; pressing a cell moves the current point along that + variable (shown as a breadcrumb), and can open the tuple's detail. For three or more variables + with many values, when the reader is exploring. + + Use small multiples or a chart in each cell to present an interaction you found, and coordinate + strips to let the reader explore. +* A view can hold as many charts as the analysis needs, each under its own section heading. + +### Draw it + +Charts are SVG written by hand (or HTML, for tables and heatmaps), drawn with the kit's parts. A +chart library is fine where it saves work (visx's scales and shapes, say), as long as what it +draws follows this guide: its own colors, fonts, tooltips, and fixed sizes do not. + +* **Width.** Draw inside `ChartFrame`, which measures its width and calls its children with it; + compute pixel coordinates from that width. Never scale a fixed drawing with `viewBox` and + `width: 100%`: the text in it would shrink in a narrow pane and grow on a wide page. Heights + are fixed per row or per chart. Wide tables scroll sideways in their own `overflow-x: auto` + wrapper; small multiples wrap (`grid-template-columns: repeat(auto-fill, minmax(240px, 1fr))`). +* **Axes and grid.** A `` holds an axis line, its ticks, and their + labels (10px, dim, tabular); `` holds grid lines, drawn before the + marks. A few round ticks, starting at zero for bars. Put units in the section's note line or a + short axis title in lower case, not on every tick. +* **Marks.** Thin and plain: bars and cells with 2px corners, starting at their baseline; dots of + 3 to 4px radius; lines 1.5 to 2px. Direct labels (`layout.chartLabel`) on the few values that + matter, not a number on every mark. +* **Legends.** A `Legend` for two or more series, placed above the chart; none for one series + (its heading names it). A `ScaleLegend` for a sequential or diverging fill. +* **Text in SVG** inherits the viewer's sans; size and color it with the `layout` classes or the + type tokens. Align with SVG attributes (`textAnchor`, `dominantBaseline`). + +## Interaction + +* **Links.** Anything that stands for an entity leads to it. In text and tables, write an + `EntityLink` (a chip the viewer draws: pressing opens the entity, hovering shows its preview). + The viewer labels the chip with the entity's type and a short ID, so put a readable name + beside it (a row's guest, a model's name), as the viewer's own tables put a key beside a value. + A chart's mark cannot be a chip, so a mark standing for an entity takes + `const links = useEntityLinks()` from `fxtr-view`: pressing it calls `links.open(id)`, which + does what pressing the entity's link does, and it carries `data-open={links.isOpen(id)}`, which + `layout.mark` (SVG) and `layout.cellMark` (HTML) draw as an ink ring while a pane shows the + entity, as the viewer lights the entity's link. A mark standing for many entities (the bar of a + mean) opens a `DetailBox` under the chart listing them as links. +* **Pressable marks.** Give every mark a reader can press `pressable(onPress, label)`, which makes + it focusable and pressable with Enter and Space and names it for screen readers ("Ada: 12 + characters"), and a `layout.mark` or `layout.cellMark` class for its hover, focus, and state + rings. +* **Hover.** Show a tooltip through `ChartFrame`'s `tooltip` prop, set on pointer enter and on + focus, cleared on leave and blur: the mark's name in bold, its value, and a line of context. + Nothing in a tooltip can be pressed, so links go in the chart or under it. Never use SVG + `` or the `title` attribute for this. +* **Picks within the view.** A reader's pick that stays inside the view (a series focused, a + point whose detail shows under the chart) is drawn with `data-picked="true"` (the same ink + ring), the rest `layout.faded`; its detail goes in a `DetailBox`. +* **Controls.** `Segmented` for two to five options, `Select` for more, labeled in lower case, + on the section heading's line at its right (as in the example below) or in one row above the + charts they control. No unstyled native controls. +* **Kept state.** What a reader chose (a tab, an order, a pick, a filter) lives in the panel's + storage, through `useStoredState`, so it survives the view remounting + ([Panel state](../reference/renderers.mdx#panel-state)). Hover and focus stay in `useState`. +* **Loading and states.** `unwrap` what the view cannot draw without; load many entities with + `LoadEach` and say how many are still loading; show a run in progress or failed with + `InvocationOutcome`; show an empty result with a `layout.empty` line ("no greetings yet"). + Handle every count of invocations a panel is given. + +## Writing + +* A view opens with its kind (`Kind` or `InvocationHead`), then the findings. No hero, no + headline question, no introductory essay, no numbered figures. +* Each chart has a lower-case section heading and at most one dim line under it saying what is + plotted, in what units, and how to read it (`layout.muted`). Join inline facts with " · ". +* Names come from the data (the model's name, the step's name), not from marketing copy. +* No emoji, and no Unicode symbols as icons (✕, ◆, ↗): draw icons in SVG with `currentColor`. + +## Example: a heatmap + +A report's scores as a model × scenario table, each cell filled by its score and opening the +transcript behind it, with the reader's choice of metric kept in storage: + +```tsx views/src/scores.tsx +import { unwrap, useEntity, useEntityLinks, type EntityData, type EntityID, type EntityPanelProps } from "fxtr-view"; +import { Kind, MISSING, ScaleLegend, Select, formatNumber, layout, pressable, sequentialFill, sequentialInk, useStoredState } from "fxtr-view-kit"; +import styles from "./scores.module.css"; + +interface ScoreData extends EntityData { + readonly $type: "my_project.Score"; + readonly model: string; + readonly scenario: string; + readonly accuracy: number; + readonly refusal: number; + readonly transcript: EntityID; +} + +interface ReportData extends EntityData { + readonly $type: "my_project.Report"; + readonly scores: readonly ScoreData[]; +} + +type Metric = "accuracy" | "refusal"; + +export function ReportPanel({ id, storage }: EntityPanelProps) { + const report = unwrap(useEntity(id as EntityID<ReportData>)); + const links = useEntityLinks(); + const [metric, setMetric] = useStoredState<Metric>(storage, "metric", "accuracy", (value): value is Metric => + value === "accuracy" || value === "refusal", + ); + const models = [...new Set(report.scores.map((s) => s.model))]; + const scenarios = [...new Set(report.scores.map((s) => s.scenario))]; + const cell = (model: string, scenario: string) => report.scores.find((s) => s.model === model && s.scenario === scenario); + return ( + <div className={layout.panel}> + <div className={layout.kindLine}><Kind>Report</Kind></div> + <div className={styles.head}> + <h3 className={layout.subtitle}>scores by model and scenario</h3> + <Select label="metric" value={metric} onChange={setMetric} + options={[{ value: "accuracy", label: "accuracy" }, { value: "refusal", label: "refusal rate" }]} /> + </div> + <p className={layout.muted}>Share of calls, 0 to 1. Press a cell to open its transcript.</p> + <ScaleLegend scale="sequential" low="0" high="1" /> + <div className={styles.scroll}> + <table className={layout.table}> + <thead> + <tr><th>model</th>{scenarios.map((s) => <th key={s} className={layout.num}>{s}</th>)}</tr> + </thead> + <tbody> + {models.map((model) => ( + <tr key={model}> + <td className={styles.key}>{model}</td> + {scenarios.map((scenario) => { + const score = cell(model, scenario); + if (!score) return <td key={scenario} className={layout.num}>{MISSING}</td>; + const value = score[metric]; + return ( + <td key={scenario} className={`${layout.num} ${layout.cellMark}`} + style={{ background: sequentialFill(value), color: sequentialInk(value) }} + data-open={links.isOpen(score.transcript)} + {...pressable(() => links.open(score.transcript), `${model} on ${scenario}: ${formatNumber(value, 2)}`)}> + {formatNumber(value, 2)} + </td> + ); + })} + </tr> + ))} + </tbody> + </table> + </div> + </div> + ); +} +``` + +```css views/src/scores.module.css +.head { display: flex; align-items: center; justify-content: space-between; gap: 12px; margin-top: 14px; } +.head > h3 { margin: 0; } +.scroll { overflow-x: auto; } +.key { font-family: var(--font-mono); font-size: var(--text-sm); color: var(--fg-dim); } +``` + +## Check it + +Save, and let the running `uv run fxtr view` rebuild the bundle (the open viewer updates in place), +or build it with `uv run fxtr views build`, which prints why a build failed +([Building and reloading](../reference/renderers.mdx#building-and-reloading)). Then search the +package for what the viewer never uses; each of these should find nothing: + +```bash +grep -rnE '#[0-9a-fA-F]{3,8}\b|rgba?\(|hsla?\(' views/src +grep -rnE 'font-family: *[^v ]|font-size: *[0-9]|color-scheme|!important|<title>' views/src +``` + +The rest is a look at the view in the viewer. The usual things to catch are a theme it reads badly +in (the toggle is at the foot of the viewer's left rail) and a width it breaks at: drag a pane +narrow, or close every pane to see the card previews. diff --git a/plugins/fxtr/skills/fxtr/references/guides/importing-external-datasets.mdx b/plugins/fxtr/skills/fxtr/references/guides/importing-external-datasets.mdx new file mode 100644 index 0000000..417829e --- /dev/null +++ b/plugins/fxtr/skills/fxtr/references/guides/importing-external-datasets.mdx @@ -0,0 +1,127 @@ +--- +title: "Importing external datasets" +description: "Bring external datasets into fxtr as typed arrays for processing in workflows." +--- + +Import an external dataset into fxtr as a typed array, then pass it to a workflow to process it. +The job records the data it uses, so you can inspect which inputs produced each result and run +other analyses against the same dataset. + +There are two recommended ways to make an imported dataset available to jobs: + +* **As a workflow input built in the launcher.** The launcher reads and validates the file and + passes a typed array to the job. This suits a fixed file. +* **As a project array.** A named dataset kept in the database, which jobs run against as + snapshots. This suits a dataset you curate over time. + +Configuration fixed in the experiment's code goes in as a **source array**, added inside the +workflow with `context.add_source_array(...)`. Credentials belong in the environment, never in +inputs. In every case, the launcher is the only code that reads files: steps and workflows are +hermetic (see +[Hermeticity](../concepts/workflows-and-steps.mdx#hermeticity-and-handling-external-state)). + +## External datasets + +Read and validate external data in the launcher, convert it to a typed array, and pass it to the +workflow. For a CSV with columns `case_id,prompt,expected`: + +```python +import csv +from pathlib import Path +from typing import TypedDict + +from fxtr.entity_defns.array import Array + + +class Case(TypedDict): + prompt: str + expected: str + + +def load_cases(path: Path): + with path.open(newline="", encoding="utf-8") as file: + return Array.from_records(csv.DictReader(file), type=("[case_id: str]", Case)) +``` + +Then, inside an [async launcher](running-and-viewing-jobs.mdx#launch-from-python): + +```python +cases = load_cases(Path("cases.csv")) +result = await run_job(client, experiment, {"cases": cases}, root="cases-v1") +``` + +Submitting stores the input array with the job, so the job keeps the exact data it ran on even if +the CSV changes later. The workflow receives a `[case_id: str]` handle to map steps over, and can +pass individual fields to steps with `cases.field("prompt")` and `cases.field("expected")`. + +When loading a dataset: + +* Key rows by stable IDs from the source, not by row number, so that reordering or adding rows + doesn't change which cached result belongs to which row. +* Convert values to the declared types yourself. `Array.from_records` validates the schema and + rejects duplicate keys, but it doesn't convert numbers, booleans, or missing CSV values. +* Check dataset-specific requirements, such as nonempty prompts, in the loader. +* Read files only in the launcher. Steps and workflows must be hermetic: they read no files and + depend on nothing outside their inputs, so that the job's record is complete and anyone can + resume it. See [Hermeticity](../concepts/workflows-and-steps.mdx#hermeticity-and-handling-external-state). + +## Configuration + +Anything that affects a result should reach its step as an input: data, model IDs, prompts, and +sampling settings. Changing one of them then changes what the step receives, so a result computed +from the old settings isn't reused (see [Managing the step cache](../concepts/managing-the-step-cache.mdx)). + +For settings fixed in the experiment's code, add a source array inside the workflow: + +```python +judge_rubric = context.add_source_array("judge_rubric", Array.scalar(RUBRIC, str)) +``` + +This makes the rubric an input of every step that receives it, so it is part of those steps' +cache fingerprints and a changed rubric is never matched with an old result. A step that read +`RUBRIC` directly from the module would have the same fingerprint whatever the constant said, +which is the situation to avoid. + +Settings used together often belong in one entity, such as a judge configuration holding the +model ID, rubric, and sampling parameters. Pass a reference to it as an input, so every result +points at the exact settings that produced it. See +[Calling language models](../../../behaviors/references/guides/language-models.mdx). + +## Entity-valued inputs + +To pass entities, store each one with the client first, then build an array from the returned +references: + +```python +async with open_local_client(__file__) as client: + intro = await client.store(Document(title="Intro", text="...")) + documents = Array.from_items([(("intro",), intro)], ("[doc: str]", BoundID[Document])) + await run_job(client, experiment, {"documents": documents}, root="docs-v1") +``` + +The only way to get an entity back is its ID, so something persisted must reference it: here the +job's input array does. An entity stored in a launcher and referenced by nothing can't be found +again, so always pass the references on as a job input or write them into a project array. + +## Project arrays + +A **project array** is a named, editable, versioned array stored in the project's database. It +suits a dataset you curate over time and launch many jobs against. A job always runs against an +immutable **snapshot**, so editing the array later doesn't change what earlier jobs used: + +```python +await client.create_array("word_count_prompts", ("[prompt: str]", str)) +await client.replace_array("word_count_prompts", prompts) +snapshot = await client.snapshot_array("word_count_prompts") +result = await run_job(client, experiment, {"prompts": snapshot.array}, root="word-count-v1") +``` + +`replace_array` sets the complete contents; `upsert_rows` adds or updates the given rows and keeps +the rest. From the command line, pass a snapshot as an input with `@`: + +```bash +uv run fxtr run word_count.experiment --input prompts=@word_count_prompts +``` + +Project arrays are optional. When you don't need a named, editable dataset, pass an ordinary +`Array` to `run_job` directly. diff --git a/plugins/fxtr/skills/fxtr/references/guides/running-and-viewing-jobs.mdx b/plugins/fxtr/skills/fxtr/references/guides/running-and-viewing-jobs.mdx new file mode 100644 index 0000000..7a883dd --- /dev/null +++ b/plugins/fxtr/skills/fxtr/references/guides/running-and-viewing-jobs.mdx @@ -0,0 +1,240 @@ +--- +title: "Running and viewing jobs" +description: "Launch jobs from Python or the command line, watch them in the viewer, read their results, and resume, inspect, or share them." +--- + +Every job runs on the project's Postgres database, configured by `fxtr init` (see the +[installation guide](../installation.mdx)). Launchers talk to it through the **project client**, which +stores entities and manages jobs, project arrays, and the cache, and the project's **viewer** shows +its jobs as they run and after they finish. + +## Before you launch + +* Commit first. A launch records the commit it runs and checks that the virtual environment + matches `uv.lock`, so it refuses uncommitted changes. Pass `allow_staged=True` (or + `--allow-staged` on the command line) to include staged changes; `run_job`, `launch_cli`, + `client.submit`, and `client.resume` all take it. `verify_environment=False` on + `open_local_client` or `launch_cli` (`--no-verify` on the command line) skips the check of the + environment, but never the pinning of the commit. Each attempt at a job records the commit it + ran and whether the environment was checked. +* Register your modules. Resuming a job, `fxtr run`, and the viewer import the modules in + `[tool.fxtr] modules`, not your launcher, so every step and workflow a job uses must be defined + in one of them. +* Use a scratch schema for trials. To keep trial jobs and cache entries apart from real + results, set `schema` in `fxtr.local.toml` to another name while experimenting (or run + `fxtr init` again with the same `--database-url`, another `--schema`, and `--force`), and set it + back for real runs. +* Start the viewer and leave it running, so each launch opens its job in it (see + [Watch jobs in the viewer](#watch-jobs-in-the-viewer)). + +## Launch from Python + +The simplest launcher calls `launch_cli`, which opens the client, submits the job, opens it in the +project's running viewer, waits, and reports the result: + +```python launch.py +from fxtr.entity_defns.array import Array +from fxtr.project.running import launch_cli +from my_project.experiment import experiment + +if __name__ == "__main__": + prompts = Array.from_items( + [(("greeting",), "Hello there"), (("question",), "Why?")], + ("[prompt: str]", str), + ) + launch_cli(experiment, {"prompts": prompts}, anchor=__file__, max_concurrency=8) +``` + +`anchor=__file__` tells fxtr which project the launcher belongs to. `max_concurrency` caps how +many steps run at once; workflow functions don't count against it. If the job doesn't succeed, +`launch_cli` exits with status 1. + +For more control, such as choosing the `root`, loading datasets, or reading results, write an +async launcher with `open_local_client` and `run_job`: + +```python launch.py +import anyio + +from fxtr.entity_defns.array import Array +from fxtr.project.running import open_local_client, run_job +from my_project.experiment import experiment + + +async def main() -> None: + prompts = Array.from_items( + [(("greeting",), "Hello there"), (("question",), "Why?")], + ("[prompt: str]", str), + ) + async with open_local_client(__file__, max_concurrency=8) as client: + result = await run_job(client, experiment, {"prompts": prompts}, root="word-count-v1") + print(list(result.items())) + + +if __name__ == "__main__": + anyio.run(main) +``` + +* `run_job` opens the job in a new tab of the project's running viewer (`uv run fxtr view`) and + prints its link before waiting, then the job's status and result; with no viewer running, it + prints how to start one. `open_browser=False`, or `FXTR_NO_BROWSER=1` in the environment, only + prints the link. It raises `JobNotFinishedError` if the job doesn't succeed. +* `root` is the root workflow's pathname, and so the start of every cache address in the job. + See [Managing the step cache](../concepts/managing-the-step-cache.mdx). +* Use `handle = await client.submit(...)` when you need the job's ID or want to control the + waiting yourself; `await wait_for_job(client, handle)` provides the same reporting as + `run_job`. +* In a notebook, call `await main()` instead of `anyio.run(main)`. + +## Watch jobs in the viewer + +Start the viewer from the project, in a terminal of its own, and leave it running: + +```bash +uv run fxtr view +``` + +It selects a local port and opens the viewer in your browser; running it again while the +project's viewer runs opens that one. The [CLI reference](../reference/cli.mdx#viewer) has its +options. Each job you launch opens in a new tab of it, at the link the launcher prints, such as +`http://127.0.0.1:8000/#job=<job id>`, and you can pick any of the project's jobs from the menu +above the minimap. A job that's still running updates live. + +For one job, the viewer shows: + +* **The tree of workflow invocations.** The root workflow at the top, with each child workflow + one level down. A mapped child appears once per key. +* **A card for every step and child workflow** a workflow scheduled, showing its docstring, how it + was mapped, its progress, and its log lines. A card's code chip, which names the commit the job + ran, opens the function's file as the job ran it in a pane, whose **Open** button opens it in your editor (choose which from its + menu). +* **The arrays each operation received and produced**, as tables you can page through and slice + by dimension. A value that references an entity is a link. +* **Entity details**, opened from any link, in a pane you can keep alongside the trace. Hovering a + link shows a preview. +* **Your project's README**, linked from the side rail. + +The first sentence of each step's and workflow's docstring is what appears on its card, so write +those as short imperative sentences that tell a reader what the operation does, such as "Compute +mean scores across judges." Leave out edge cases and what the step is mapped over, which the card +already shows. For progress worth reading while a step runs, write log lines with +`context.log(...)`. + +All of this is drawn by the viewer's default renderers until the project adds its own, which can +also give the root workflow a results overview: a view of the run that presents its results. See +[Visualization and custom renderers](../concepts/visualization-and-custom-renderers.mdx). + +## Read results + +`run_job` returns the root workflow's result as an `Array`. Entity values in it are bare IDs, so +load them with their type. + +Have the root workflow return one **report** entity that references every result you want to +present. The launcher then gets the whole run from one result, and a +[results overview](../concepts/visualization-and-custom-renderers.mdx#results-overviews) in the +viewer can read everything from it: + +```python +@dataclass(frozen=True) +class Report(DataclassEntity, fxtr_type="org.example.qa.Report.v1"): + scores: BoundID[Array[float]] + summary: BoundID[Array[float]] + + +@workflow(name="qa.evaluate") +async def evaluate(context: WorkflowContext, ...) -> Report: + """Score every answer and summarize the scores per model.""" + scores = context.run_step("scores", score, {...}, map_over=["model", "task"]) + summary = context.run_step("summary", summarize, {"scores": scores}, map_over=["model"]) + return Report( + scores=await context.store(await scores.observe()), + summary=await context.store(await summary.observe()), + ) +``` + +Observing inside the workflow is fine here: nothing computes from the report, and each array it +references is still a step output in the graph. The launcher loads the report, then each array it +references: + +```python +report = await client.load(result.item(), as_type=Report) +scores = await client.load(report.scores) # references inside a loaded entity are bound +``` + +Keep the client open while loading. Alternatively, pass an async `on_result(storage, result)` +callback to `run_job` or `launch_cli`; it replaces the default result printing and runs while +storage is open. Leaving the client's `async with` block waits for the jobs it started; leaving it +with an exception cancels them. + +## Launch from the command line + +`fxtr run` launches a workflow by its registered name: + +```bash +uv run fxtr run word_count.experiment --input prompts=@word_count_prompts --root word-count-v1 +``` + +Inputs are JSON scalars, or `@NAME` for a snapshot of a +[project array](importing-external-datasets.mdx#project-arrays). Anything richer needs a launcher script. + +## Resume a stopped job + +A job stops when it can't make further progress: a step raised, a step is in cache conflict, or +the job was cancelled. Resume it with its ID, which the launcher prints: + +```bash +uv run fxtr resume 3fa3df92-2edb-4088-9e1f-a99d28b7ec77 +``` + +Resuming uses the job's stored inputs and your registered modules, records the current commit, +reuses every result the job already has, and retries unfinished work, including steps that +raised. From Python: + +```python +import uuid + +from fxtr.project.client import JobID +from fxtr.project.running import wait_for_job + +handle = await client.resume(JobID(uuid.UUID("3fa3df92-2edb-4088-9e1f-a99d28b7ec77"))) +result = await wait_for_job(client, handle) +``` + +Add `--overwrite-cache-conflicts` (`overwrite_cache_conflicts=True` in Python) to rerun the +steps in cache conflict and replace their old results (see +[Resolving a job's conflicts](../concepts/managing-the-step-cache.mdx#resolving-a-jobs-conflicts)). A resume +continues the same job, so it can't change what the job already did; to change the experiment, +launch a new job, which reuses every unchanged step's result. + +Resume refuses a job that already succeeded, and a job another process is running. Once that +process has been silent for `[tool.fxtr] stale_after` seconds (30 by default), fxtr treats the job +as abandoned, and resume takes it over. + +## Inspect and cancel jobs + +| Command | Python | Purpose | +| ------------------------------------ | ------------------------------------------------------ | -------------------------------------------- | +| `fxtr jobs list` | `await client.list_jobs()` | Every job, newest first. | +| `fxtr jobs status JOB` | `await handle.status()` | Where a job is, and why it stopped. | +| `fxtr jobs cancel JOB` | `await handle.cancel()` | Ask the process running a job to stop it. | +| `fxtr cache list` | `await client.list_cache_entries(prefix=...)` | The cache's entries. | +| `fxtr cache list --conflicts-in JOB` | `await client.list_cache_entries(conflicts_in=job_id)` | The entries a job is in cache conflict with. | + +In Python, `handle = await client.job(job_id)` attaches to an existing job without starting it. + +## Share a job with another project + +A finished job can be written to a file and imported into another copy of the project, for +example a collaborator's, with its results and everything they reference: + +```bash +uv run fxtr jobs export 3fa3df92-2edb-4088-9e1f-a99d28b7ec77 results/qa-v1.car +uv run fxtr jobs import results/qa-v1.car --adopt-cache # in the other project +``` + +The record carries, for every result, the commit that produced it, and the importing project's +repository must contain those commits, so share the code as well as the file. An imported job +appears in the viewer and can be resumed like any other. By default its results stay out of the +step cache; `--adopt-cache` puts them in, so new jobs in the importing project reuse them (see +[Reusing another job's results](../concepts/managing-the-step-cache.mdx#reusing-another-jobs-results)). +To refresh a job that was imported earlier from a later record of it, import with `--update`. +From Python, the client has `export_job(job_id, path)` and `import_job(path, update=...)`. diff --git a/plugins/fxtr/skills/fxtr/references/index.mdx b/plugins/fxtr/skills/fxtr/references/index.mdx new file mode 100644 index 0000000..c46b31f --- /dev/null +++ b/plugins/fxtr/skills/fxtr/references/index.mdx @@ -0,0 +1,54 @@ +--- +title: "What is fxtr?" +description: "fxtr is an experiment orchestration system for verifiable AI-powered analyses" +--- + +fxtr (pronounced "fixture") is a framework for running experiments and data analyses with AI systems. It strives to give users (and their coding agents) full control over the core experiment design, including the *experiment logic*, *data types*, *input data*, and *presentation of results*, while solving complementary organizational and operational problems such as managing data storage, caching expensive steps, ensuring reproducibility, verifying what code was actually run, and sharing results. + +In short: fxtr tries to preserve the full power of coding agents and avoid restricting your experiment design, but provide enough scaffolding so that your analyses are systematic, organized, and reproducible and do not devolve into slop. + +<CardGroup cols={3}> + <Card title="Installation" icon="download" href="installation.mdx"> + Install fxtr and set up your project and database. + </Card> + + <Card title="Your first experiment" icon="flask" href="first-experiment.mdx"> + Write, launch, and rerun a small experiment. + </Card> + + <Card title="Core concepts" icon="book" href="concepts/overview.mdx"> + Projects, steps, workflows, jobs, arrays, and entities. + </Card> +</CardGroup> + +## How it Works + +1. **Define your experiment as simple, readable, environment-independent code.** Each experiment or analysis is expressed as a Python package containing one or more *workflows*, which schedule expensive or nondeterministic *steps* into a graph of dependent computations. This experiment definition serves as an ideal "spec" for what you want to compute, without worrying about data storage, concurrency, or looking up previously-computed results. The dataflow and data types belong to you and are highly customizable. +2. **Let the system drive it to completion.** fxtr traces and durably executes your experiment code, stores the data, and gives you live observability into what is running and what is scheduled. If your jobs crash or get interrupted, you can seamlessly resume them, without recomputing expensive results. +3. **Iterate freely.** Want to increase the range of your sweeps, or add a new analysis step to your experiment? Simply edit the workflow code in place and launch a new job. fxtr automatically re-uses previously computed results across jobs, as long as they have the same *name* (or *cache address*) and receive the same *inputs*. +4. **Customize data visualizations and reports.** Your coding agent can register custom renderers for your experiment workflows and data types, and have them directly integrated into the viewer. +5. (Under construction) **Share your results.** You can share your code as a git repo to enable others to reproduce your experiment, and export and import jobs to share the data from your analyses. Soon, we plan to make it possible to publish experiments and share them via a hosted platform. + +## Components of a `fxtr` Project + +fxtr experiments are organized into *projects*, which have two components. + +* The **project repo** is a Git repo that defines how to execute and visualize your experiment or analysis. It contains a few pieces: + * **Workflow functions**: Python functions that define the overall structure of your experiment, configure sweeps, and schedule individual nondeterministic steps into a directed acyclic graph. Workflow functions are expected to be deterministic, and will be re-run each time a job is started or resumed. + * **Step functions**: Python functions that execute nondeterministic or expensive work, such as querying a language model. fxtr executes each of these once and then caches the results. Later runs re-use existing step results. To ensure reproducibility, steps should be environment-independent: you should get equivalent results regardless of which computer you run the steps on. (So, querying a language model API is OK, but reading or writing files on a local filesystem is generally not.) + * **Entity types**: Data types used by your experiment. Entities can point to other entities, and are used to save data like transcripts, judgements, and experiment configurations. + * **Renderers**: React components that define the visual presentation of a workflow, step, or entity. These renderers are registered with the system and can be used to visualize and inspect the results generated by the experiment. +* The **experiment database** is a database that holds the data and results related to a project. It includes: + * **Job outputs and experiment state**: Each execution of a workflow in a project is persisted to the database, including all steps and workflows that it ran. Job records can be exported and imported to share your results. + * **The step result cache**: The result of each successful step execution is cached, and stored under a *cache address*. Future jobs will re-use results from the cache, avoiding duplicate work. The cache can also be cleared manually to avoid conflicts or regenerate results from scratch. + * **Project arrays**: Arrays of data that you directly store into the database from an external source. These arrays can then be passed as inputs to workflow functions when starting a job. + +## Getting Started + +fxtr pairs with [**behaviors**](../../behaviors/references/index.mdx), a library for calling language models from fxtr steps. It +provides model adapters for OpenAI and Anthropic, multi-turn conversations, retries, and +storable conversation records. + +To get started, follow the [installation guide](installation.mdx) to install fxtr and set up your project and database. +The guide also covers fxtr's [agent skills](installation.mdx#install-the-agent-skills), which let +Claude Code set up projects and write and run experiments for you. diff --git a/plugins/fxtr/skills/fxtr/references/installation.mdx b/plugins/fxtr/skills/fxtr/references/installation.mdx new file mode 100644 index 0000000..a85ca14 --- /dev/null +++ b/plugins/fxtr/skills/fxtr/references/installation.mdx @@ -0,0 +1,174 @@ +--- +title: "Installation" +description: "Install fxtr and get started with an example project" +--- + +<Tip> + fxtr ships skills for Claude Code and Codex. + Once they're [installed](#install-the-agent-skills), you can ask your agent to create a + project, connect it to your database, and write, launch, and fix experiments, then review its + work in the viewer. +</Tip> + +<Steps> + <Step title="Prerequisites" id="prerequisites"> + Before continuing, make sure you have the following tools installed on your machine: + + | Tool | Install | + | ------------------------------------------------ | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | + | [uv](https://docs.astral.sh/uv/) 0.12.2 or newer | `curl -LsSf https://astral.sh/uv/install.sh \| sh`, or `uv self update` to upgrade. Other methods: [uv installation](https://docs.astral.sh/uv/getting-started/installation/). | + | git | Usually preinstalled. Otherwise see [Installing Git](https://git-scm.com/book/en/v2/Getting-Started-Installing-Git). jj also works. | + | [Docker](https://www.docker.com/) | On macOS, [Docker Desktop](https://docs.docker.com/desktop/setup/install/mac-install/). On Linux, [Docker Engine](https://docs.docker.com/engine/install/) for your distribution. | + | [Node](https://nodejs.org/) 22 or newer | The [Node download page](https://nodejs.org/en/download) gives commands for each system and version manager. | + | [pnpm](https://pnpm.io/) 10 | `npm install -g pnpm@10`, once Node is installed. Other methods: [pnpm installation](https://pnpm.io/installation). | + </Step> + + <Step title="Start Postgres" id="start-postgres"> + fxtr keeps every job, result, and cache entry in Postgres. To run a local server in a container: + + ```bash + docker run -d --name fxtr-postgres -p 55433:5432 \ + -e POSTGRES_USER=fxtr -e POSTGRES_PASSWORD=fxtr -e POSTGRES_DB=fxtr postgres:16 + ``` + + It listens on port `55433`, with user, password, and database all named `fxtr`, so its URL is + `postgresql+asyncpg://fxtr:fxtr@localhost:55433/fxtr`. + </Step> + + <Step title="Install the agent skills" id="install-the-agent-skills"> + Install the fxtr plugin in your coding agent: + + <Tabs> + <Tab title="Claude Code"> + In Claude Code: + + ```text + /plugin marketplace add TransluceAI/claude-code-plugins + /plugin install fxtr@transluce-plugins + ``` + </Tab> + + <Tab title="Codex"> + From a terminal: + + ```bash + codex plugin marketplace add TransluceAI/codex-plugins + codex plugin add fxtr@transluce-plugins + ``` + </Tab> + </Tabs> + + The skills teach your agent how to write experiment logic, import data, operate the runner and viewer, manage their state, and much more. Just ask for what you want! + </Step> +</Steps> + +## Run an example + +To check that it works, create and run a complete greeting example. The +`--example` flag includes the experiment code and its launcher, so you don't write any code for +this check. (With the skills installed, you can also ask your agent to create the fxtr example +project and run it.) + +Without `--example`, `fxtr new` creates an empty project with no experiment or launcher. +[Your first experiment](first-experiment.mdx) walks through that path from project creation +to writing and running your own code. + +<Steps> + <Step title="Create the example project"> + ```bash + mkdir ~/fxtr + uvx fxtr new ~/fxtr/fxtr-example --example + ``` + + This writes a uv project that depends on fxtr, with a complete greeting experiment in + `src/fxtr_example/experiment.py`, a `launch.py` that runs it, and greeting renderers for + the viewer. It also runs `git init` and prints the next steps. + + Install the example's Python and renderer dependencies, then build its renderers: + + ```bash + cd ~/fxtr/fxtr-example + uv sync + (cd views && pnpm install) + uv run fxtr views build + ``` + </Step> + + <Step title="Connect it to the database"> + From the new project: + + ```bash + uv run fxtr init --database-url postgresql+asyncpg://fxtr:fxtr@localhost:55433/fxtr + ``` + + This writes `fxtr.local.toml`, which records where this machine reaches the database and is + never committed, and creates the project's tables. + </Step> + + <Step title="Start the viewer"> + In a second terminal, from the project: + + ```bash + uv run fxtr view + ``` + + This serves the experiment viewer at `http://127.0.0.1:8000/` and opens it in your browser. + </Step> + + <Step title="Commit and launch the example"> + Every launch records the commit it runs, so commit first: + + ```bash + git add . && git commit -m "A new fxtr project" + uv run python launch.py + ``` + + The included code greets three guests: + + ```text + job f0d6b1ac-a781-4b93-a571-edef402ad131: running + watch it at http://127.0.0.1:8000/#job=f0d6b1ac-a781-4b93-a571-edef402ad131 + job f0d6b1ac-a781-4b93-a571-edef402ad131: succeeded + result: bafyreid7ea3tcab7lyxpgpyisaiwphsyqygdjnynpupdgqpk3tc2iixxsi + Hello, Ada! + Hello, Alan! + Hello, Grace! + ``` + + The job opens in a new tab of the viewer, showing its graph: the guest list, the `greet` step + mapped over it, and each greeting it produced. + </Step> +</Steps> + +## What's in the example project + +```text +pyproject.toml the project, its dependency on fxtr, and [tool.fxtr] +README.md how to run it +launch.py runs the experiment as a job +src/fxtr_example/ + experiment.py an entity type, a step, and a workflow: an example to replace +views/ renderers that customize how the viewer shows your data + vendor/fxtr/ the fxtr TypeScript packages the renderers build against +``` + +`[tool.fxtr] modules` in `pyproject.toml` lists the modules that define your steps and +workflows. When you add an experiment module, add it there too. +[Project configuration](reference/project-configuration.mdx) describes each setting and the +files `fxtr new` writes. + +The renderers' fxtr packages aren't on npm: fxtr carries them, and `views/vendor/fxtr/` is the +project's copy. Commit it. After upgrading fxtr, refresh it with `uv run fxtr views vendor`, then +run `pnpm install` in `views/` and rebuild. + +## Next steps + +<CardGroup cols={2}> + <Card title="Your first experiment" icon="flask" href="first-experiment.mdx"> + Write an experiment of your own, step by step. + </Card> + + <Card title="Core concepts" icon="book" href="concepts/overview.mdx"> + How projects, steps, workflows, jobs, and arrays fit together. + </Card> +</CardGroup> diff --git a/plugins/fxtr/skills/fxtr/references/reference/array-queries.mdx b/plugins/fxtr/skills/fxtr/references/reference/array-queries.mdx new file mode 100644 index 0000000..d1251b6 --- /dev/null +++ b/plugins/fxtr/skills/fxtr/references/reference/array-queries.mdx @@ -0,0 +1,513 @@ +--- +title: "Array queries" +description: "Filter, reshape, aggregate, and join arrays with the query builder or SQL, in a workflow or in a step." +--- + +This page is the reference for querying arrays: filtering, reshaping, aggregating, and joining +them, in a workflow or in a step. Write a query with Python methods such as `filter()` and +`aggregate()`, called the query builder, or with SQL through `fxtr.sql(...)`. +[Querying, filtering, and aggregating arrays](../concepts/arrays-and-parallel-computations.mdx#querying-filtering-and-aggregating-arrays) +introduces both. + +Writing a query does not run it; `collect()` does. Over arrays (`Array`), `collect()` runs the +query and returns the result. In a workflow, over handles (`ArrayHandle`), it adds the query to the +workflow's graph as a query node instead. + +## An example + +This experiment removes the runs that ended in an error, then summarizes each model's remaining +runs: how many scored at least `0.8`, and the mean score of those runs. + +```python src/my_project/evaluation.py +from typing import Annotated, TypedDict + +import fxtr +from fxtr.entity_defns.array import Array +from fxtr.entity_defns.query import value +from fxtr.experiment.handles import ArrayHandle +from fxtr.experiment.steps import StepContext, step +from fxtr.experiment.workflows import WorkflowContext, workflow + + +class Run(TypedDict): + score: float + error: str | None + + +class Summary(TypedDict): + passed: int + mean_score: float | None + + +@step(name="evaluation.summarize") +async def summarize( + context: StepContext, runs: Annotated[Array[Run], "[run: int]"] +) -> Summary: + """Count a model's passing runs and average their scores.""" + stats = await fxtr.sql( + """ + SELECT {passed: count(*), mean_score: avg(value.score)} AS value + FROM runs WHERE value.score >= 0.8 + """, + tables={"runs": runs}, + output_type=Summary, + ).collect() + return stats.item() + + +@workflow(name="evaluation.evaluate") +async def evaluate( + context: WorkflowContext, runs: Annotated[ArrayHandle[Run], "[model: str, run: int]"] +) -> Annotated[ArrayHandle[Summary], "[model: str]"]: + """Summarize each model's runs that finished without an error.""" + finished = runs.filter(value().field("error").is_null()).collect() + return context.run_step("summarize", summarize, {"runs": finished}, map_over=["model"]) +``` + +Launch it with a `[model: str, run: int]` array of runs as its input (see +[Launch from Python](../guides/running-and-viewing-jobs.mdx#launch-from-python)). The pieces: + +* **In the workflow**, `runs.filter(...)` builds a query over the handle `runs`; + `value().field("error")` reads each run's `error` field (see [Build a query](#build-a-query)). + `collect()`, which is not awaited, adds the query to the workflow's graph as one query node and + returns the node's handle, `finished`. The rows never reach the workflow body: the runner + computes the node, and `run_step` receives its handle. +* **In the step**, `runs` arrives as an `Array`, so `fxtr.sql(...)` builds a query over it; in + SQL, `value.score` is each row's `score` field (see [Write SQL](#write-sql)). + `await ... .collect()` runs the query in the step's process and returns an `Array`. The result + is one value with no dimensions, which `item()` reads as a `Summary`. +* **Mapping** over `model` calls `summarize` once for each model that has a row in `finished`. A + model whose runs all ended in an error has no rows there, so the result has no row for it (see + [Mapping logic over arrays](../concepts/arrays-and-parallel-computations.mdx#mapping-logic-over-arrays)). + +In the viewer, a query appears inside the argument of the step that reads it: `summarize`'s +`runs` argument reads `query(runs)`. + +## Write SQL + +`fxtr.sql(...)` builds a query from one read-only DuckDB `SELECT`, so you write DuckDB's dialect of +SQL, including CTEs, joins, window functions, and list functions. Like a builder query, it runs or +is recorded when you call `collect()`. + +The examples here and in [Build a query](#build-a-query) use two arrays: `runs`, which holds each +model's scored runs, and `models`, which holds each model's owner. The comment under each example +shows what `collect()` returns: the result's type, then its rows as `(dimension keys): value`, +with `...` for rows or fields left out. The same queries over handles holding these rows give the +same results once collected. + +```python +from typing import TypedDict + +from fxtr.entity_defns.array import Array + + +class Run(TypedDict): + score: float + error: str | None + + +class Model(TypedDict): + owner: str + + +runs = Array.from_records( + [ + {"model": "a", "run": 1, "score": 0.9, "error": None}, + {"model": "a", "run": 2, "score": 0.4, "error": "timeout"}, + {"model": "b", "run": 1, "score": 0.7, "error": None}, + {"model": "b", "run": 2, "score": 0.95, "error": None}, + ], + ("[model: str, run: int]", Run), +) +models = Array.from_records( + [{"model": "a", "owner": "alice"}, {"model": "b", "owner": "bob"}], + ("[model: str]", Model), +) +``` + +In SQL, each table has two columns: `dims`, a struct of the row's dimension keys, and `value`, the +row's value. So `runs.dims.model` is a run's model, and `runs.value.score` is its score. + +Find the best run for each model. `$per_model` is a parameter, bound to its value rather than +pasted into the SQL: + +```python +import fxtr + +best = fxtr.sql( + """ + SELECT dims, value FROM runs + QUALIFY row_number() OVER (PARTITION BY dims.model ORDER BY value.score DESC, dims.run) + <= $per_model + """, + tables={"runs": runs}, + parameters={"per_model": 1}, +) +# [model: str, run: int]: Run +# ("a", 1): {score: 0.9, error: None} ("b", 2): {score: 0.95, error: None} +``` + +Report each model's number of runs, owner, and mean score, joining the two tables: + +```python +from fxtr.entity_defns.schema import optional, struct_of + +report = fxtr.sql( + """ + SELECT {model: runs.dims.model} AS dims, + {runs: count(*), owner: any_value(models.value.owner), + mean_score: avg(runs.value.score)} AS value + FROM runs JOIN models ON runs.dims.model = models.dims.model + GROUP BY runs.dims.model + """, + tables={"runs": runs, "models": models}, + output_type=("[model: str]", struct_of(runs=int, owner=str, mean_score=optional(float))), +) +# [model: str]: {runs: int, owner: str, mean_score: float | None} +# ("a",): {runs: 2, owner: "alice", mean_score: 0.65} ("b",): {runs: 2, owner: "bob", ...} +``` + +Total the scores of all runs: + +```python +total = fxtr.sql( + "SELECT coalesce(sum(value.score), 0.0) AS value FROM runs", + tables={"runs": runs}, + output_type=float, +) +# []: float +# (): 2.95 +``` + +SQL's `sum` of no rows is null, while the builder's `sum()` gives `0`. `coalesce(..., 0.0)` keeps +`total` a valid `float` after a filter that removes every row. To keep the null instead, declare +`optional(float)`. + +### Shape the result + +Return a `dims` column holding the result's dimension keys and a `value` column holding its +values. A result with no dimensions, such as `total`, needs only `value`. Dimension keys must be +non-null `int` or `str` values, unique for each row; to return repeated rows, add a +`row_number()` dimension. The result is always in dimension-key order; `ORDER BY` matters only +inside a window function and for choosing which rows a `LIMIT` keeps. + +**With one table, the result's type defaults to that table's type.** That suits SQL that only +chooses rows, such as `best`. Set `output_type` when the SQL changes the dimensions or the value's +schema, as `total` does, and whenever it reads more than one table, as `report` does. Without it, +the query fails when it runs, with an error saying the columns do not match. + +An integer column widens to a declared `float`: `SELECT 5` declared `float` gives `5.0`. Any other +mismatch between a column and its declared type is an error, so cast in SQL, as in +`NULL::VARCHAR`. + +### Mix SQL with the builder + +A table can be a builder query, such as `runs.filter(...)`, and an SQL result has the builder's +methods. The tables must be all arrays and `ArrayQuery`s, which gives an `ArrayQuery`, or all +handles and `ArrayHandleQuery`s from the same run of a workflow body, which gives an +`ArrayHandleQuery`. In a workflow where `runs` is a handle, this finds each model's best finished +run and records it as one query node: + +```python +from fxtr.entity_defns.query import value + +best_finished = fxtr.sql( + """ + SELECT dims, value FROM finished + QUALIFY row_number() OVER (PARTITION BY dims.model ORDER BY value.score DESC, dims.run) = 1 + """, + tables={"finished": runs.filter(value().field("error").is_null())}, +).collect() +``` + +### What SQL can read + +SQL reads only the tables in `tables` and its own `WITH` queries, named without a schema. Multiple +statements, writes, file or network access, extensions, and table functions other than `unnest`, +`range`, and `generate_series` are rejected when the query is built. + +In SQL, an entity ID is its content-derived ID as bytes. SQL can compare, copy, and return +entity IDs, and can build one from bytes. Storing the result fails if it references an entity +that is missing from storage. In a workflow, the result is stored when it is passed to a step or +a child workflow, or returned by the workflow, and that failure fails the attempt (the current run +of the workflow body). + +A list is one value, so a step can't be mapped over its elements. SQL can turn the elements into +a dimension with `unnest`, and fold a dimension back into a list with `list(...)`: see +[Turning lists into dimensions](../concepts/arrays-and-parallel-computations.mdx#turning-lists-into-dimensions). +The query builder has no `unnest`, so this needs SQL. + +## Build a query + +The examples use `runs` and `models` from [Write SQL](#write-sql); `model` and `run` are +dimensions of `runs`, not fields of its values. The expression helpers all come from +`fxtr.entity_defns.query`: `dim`, `value`, `literal`, `when`, `asc`, `desc`, and `count`. The +first example below imports the ones the examples use. + +Use `dim("model")` to read a dimension, `value()` to read the whole row value, and +`value().field("score")` to read a field of a struct value. Chain `field` for a nested struct, as +in `value().field("metrics").field("score")`; a field name is used as written, so +`field("a.b")` reads a field named `a.b`. A dimension and a field can have the same name, since +`dim` and `value().field` name which one they read. + +### Filter rows + +`filter(predicate)` keeps the rows where the predicate is true: + +```python +from fxtr.entity_defns.query import count, desc, dim, value, when + +failed = runs.filter(value().field("error").is_not_null() & (value().field("score") < 0.5)) +# [model: str, run: int]: Run +# ("a", 2): {score: 0.4, error: "timeout"} +``` + +**Write `&`, `|`, `~`, `is_in([...])`, and `is_null()`, not `and`, `or`, `not`, `in`, and +`is None`.** Python's `&` and `|` bind more tightly than comparisons, so put each comparison in +parentheses. `and`, `or`, `not`, and chained comparisons such as `a == b == 2` raise a `TypeError` +that names the operators to use. A bool-typed expression is a predicate by itself: for an array with +a bool field `passed`, write `array.filter(value().field("passed"))`. Nulls follow SQL's rule; see +[Where queries differ from Python](#where-queries-differ-from-python). + +### Reshape values + +```python +scores = runs.select(value().field("score")) +# [model: str, run: int]: float +# ("a", 1): 0.9 ("a", 2): 0.4 ("b", 1): 0.7 ("b", 2): 0.95 + +score_rows = runs.select({"score": value().field("score")}) +# [model: str, run: int]: {score: float} +# ("a", 1): {score: 0.9} ... + +model_a = runs.sel({"model": "a"}) +# [run: int]: Run +# (1,): {score: 0.9, error: None} (2,): {score: 0.4, error: "timeout"} +``` + +One expression replaces each row's value, and a mapping builds a struct from several; neither +changes the dimensions. After a `select`, `value()` refers to the new value. `runs.field("score")` +is shorthand for the first form. `sel` keeps the rows at the given dimension values and removes +those dimensions; despite the similar name, it is unrelated to `select`. +`rename({"run": "attempt"})` renames dimensions. + +### Aggregate rows + +```python +per_model = runs.aggregate( + {"runs": count(), "mean_score": value().field("score").mean()}, + group_by={"model": dim("model")}, +) +# [model: str]: {runs: int, mean_score: float | None} +# ("a",): {runs: 2, mean_score: 0.65} ("b",): {runs: 2, mean_score: 0.825} +``` + +The `group_by` names become the result's dimensions, in the order given, and group keys must be +non-null `int` or `str` values. Without `group_by`, the result has no dimensions. Bare `count()` +counts rows, and `expression.count()` counts non-null values. The other aggregations are methods +on expressions: `sum()`, `mean()`, `min()`, `max()`, `n_unique()`, `std()` and `var()` (sample +statistics), `quantile(q)` (linear interpolation), and `median()`. + +Every aggregation except `n_unique()` skips nulls; `n_unique()` counts null as one value. This +table shows what each gives when there are no rows, and when every value is null: + +| Aggregation | No rows | Only nulls | +| ----------------------------------------------------- | ------------------------- | ------------------------------ | +| `count()` | `0` | the number of rows | +| `expression.count()`, `sum()` | `0` | `0` | +| `n_unique()` | `0` | `1` (null counts as one value) | +| `mean()`, `min()`, `max()`, `median()`, `quantile(q)` | null | null | +| `std()`, `var()` | null (also for one value) | null | + +### Keep the top rows + +```python +best_two = runs.limit(2, order_by=desc(value().field("score"))) +# [model: str, run: int]: Run +# ("a", 1): {score: 0.9, error: None} ("b", 2): {score: 0.95, error: None} +``` + +`limit(n, order_by=...)` keeps the first `n` rows under one or more sort keys. Each key is an +expression, sorted ascending, or `asc(...)` or `desc(...)`. When rows tie on the sort keys, their +dimension keys break the tie, so the result is the same every time. With `ties="all"`, `limit` also +keeps every row that ties with the last row kept, so it can return more than `n` rows; it needs +`order_by`. The sort only chooses rows: the result, like every array, is in dimension-key order, +as above. + +### Join arrays + +`join` keeps the pairs of rows whose columns are equal for every pair of columns in `on` (an inner +join). The result's value holds the two rows' values under `left` and `right`, which `left_name=` +and `right_name=` rename. **The result's dimensions are the left input's, followed by the right +input's that are not join keys.** A right dimension whose name the left input already uses is +renamed with the `right_name` value and an underscore as a prefix, as in `right_run` (or `b_run` +with `right_name="b"`): + +```python +owned = runs.join(models, on=[(dim("model"), dim("model"))]) +# [model: str, run: int]: {left: Run, right: Model} +# ("a", 1): {left: {score: 0.9, error: None}, right: {owner: "alice"}} ... + +a_runs, b_runs = runs.sel({"model": "a"}), runs.sel({"model": "b"}) +paired = a_runs.join(b_runs, on=[(dim("run"), dim("run"))], left_name="a", right_name="b") +# [run: int]: {a: Run, b: Run} +# (1,): {a: {score: 0.9, ...}, b: {score: 0.7, ...}} +# (2,): {a: {score: 0.4, ...}, b: {score: 0.95, ...}} + +same_error = a_runs.join(b_runs, on=[(value().field("error"), value().field("error"))]) +# [run: int, right_run: int]: {left: Run, right: Run}, with no rows: run 1's error is null on +# both sides, and null never equals null +``` + +The dimensions decide how a step maps over the result: `paired` has one row per run, so a step +comparing the two models maps over `run`. + +`context.combine` is the other way to put a workflow's handles side by side. It lines them up +along their shared dimensions, broadcasting where one lacks a dimension, into one struct per key, +with no join condition. Use `join` to match an explicit pair of columns, such as one array's +dimension against another's field. + +### Build expressions + +Expressions support arithmetic, `cast(t)` (to `int`, `float`, `str`, `bool`, or `bytes`), +`fill_null(x)`, `is_in([...])`, `between(low, high)` (inclusive), and +`when(...).then(...).otherwise(...)`: + +```python +score = value().field("score") +computed = runs.select( + { + "percent": score * 100, + "error": value().field("error").fill_null("none"), + "chosen_model": dim("model").is_in(["a", "c"]), + "band": when(score.between(0.8, 1.0)).then("high").otherwise("other"), + } +) +# [model: str, run: int]: {percent: float, error: str, chosen_model: bool, band: str} +# ("a", 1): {percent: 90.0, error: "none", chosen_model: True, band: "high"} +# ("a", 2): {percent: 40.0, error: "timeout", chosen_model: True, band: "other"} ... +``` + +The same expressions work in `select`, `filter`, aggregations, `group_by`, and `order_by`. Use +`literal(x)` for a standalone constant, and `literal(None, optional(int))` for a typed null, with +`optional` from `fxtr.entity_defns.schema`. A query can compare, copy, join on, and write entity +IDs, but cannot look inside them: a query never loads an entity. For anything else, such as window +functions, list or string functions, or recursive CTEs, use [SQL](#write-sql). + +## Run a query now, or record it in a workflow + +The same methods work on an `Array` and on an `ArrayHandle`. Called on either, `filter`, +`select`, `field`, `sel`, `rename`, `limit`, `aggregate`, and `join` start a query: +`array.filter(p)` is short for `array.query().filter(p)`. Called on a query, they extend it. +What `collect()` does depends on what the query reads: + +* **Over arrays** (an `ArrayQuery`, in a step, a notebook, or a launcher), `await query.collect()` + runs the query immediately and returns an `Array`. The query keeps no result, so each + `collect()` runs it again; to use a result twice, keep the returned array. +* **Over handles** (an `ArrayHandleQuery`, in a workflow), `query.collect()` runs nothing and is + not awaited. It adds the query to the workflow's graph as a query node, which the runner + computes as soon as its inputs have values, and returns the node's handle. Pass the handle to a + step, query it again, or wait for its rows with `await handle.observe()` or + `await handle.item()`. Collecting the same query twice adds one node, and a query that is never + collected adds none. To reuse part of a query, collect that part and build on the handle. + +A query has no `item()`, so collect it first and read values from the result. Over arrays, +write `(await query.collect()).item()`. Over handles, write `await query.collect().item()`: +`collect()` returns the handle without waiting, and `item()` waits for the value. +**`item()` needs an array with no dimensions.** One row is not enough, so +`(await runs.limit(1).collect()).item()` raises `ArrayTypeError`, from +`fxtr.entity_defns.array`. Read a result with dimensions by key, as in `result[("a", 1)]`, or row +by row with `result.items()`. + +**Collect a query before passing it to `run_step`**, which takes handles, as `evaluate` does in +[the example](#an-example). + +**One query reads either arrays or handles, not both.** Mixing them raises a `QueryError`. To +query an array together with handles, add it to the workflow with `context.add_source_array(...)`; +to query a handle's rows as an array, `observe()` it. + +## Query in the workflow or in a step + +When a workflow needs to filter, reshape, aggregate, or join results before the next step, the +work can go in the workflow body, as a query, or in a step, which receives the arrays and may +query them itself, as `summarize` does and as reduction steps do. What differs: + +* A query in the workflow body needs no step function and adds nothing to the step cache. The + runner computes it again on a resumed attempt and in a later job. +* A step's result is cached, and a step can do what a query cannot: call a model or another + service, load an entity's contents, or run arbitrary Python. + +So use a step when the work calls a model or another service, needs an entity's contents or +arbitrary Python, is slow enough that its result should be cached, or must give one answer +however often it runs, such as a random draw or a timestamp (see +[Keep query nodes stable](#keep-query-nodes-stable)). For a transformation a query can express, +either works. + +For example, `summarize` in [the example](#an-example) can also be one query over `finished`. +`summarize` runs once per model and always returns one row, so a model whose finished runs all +score under `0.8` gets `passed: 0`. A query that filtered on the score first would have no row for +that model, so this query makes the score conditional instead. In `evaluate`, it replaces the +`run_step` call, and the workflow returns `summaries`: + +```python +from fxtr.entity_defns.query import dim, literal, value, when +from fxtr.entity_defns.schema import optional + +score = value().field("score") +passing = when(score >= 0.8).then(score).otherwise(literal(None, optional(float))) +summaries = finished.aggregate( + {"passed": passing.count(), "mean_score": passing.mean()}, group_by={"model": dim("model")} +).collect() +``` + +## Where queries differ from Python + +Queries follow SQL's rules, as DuckDB implements them: + +* **A comparison with null is null, and `filter` drops rows whose predicate is null.** With + `error = value().field("error")`, `runs.filter(error != "timeout")` returns no rows: the three + runs without an error have a null `error`, and `null != "timeout"` is null. + `runs.filter(~(error == "timeout"))` returns no rows either. To keep those three runs, write + `runs.filter(error.is_null() | (error != "timeout"))`. Likewise, `is_in([2, None])` is null, + not false, for a value other than `2`. +* **`&` and `|` evaluate both sides.** For a string field `s`, `(s != "") & (s.cast(int) > 10)` + still converts the empty strings, and fails on them. Guard the cast with an earlier `filter`, or + with `when(s != "").then(s.cast(int) > 10).otherwise(False)`. +* **Integer arithmetic is SQL's.** `//` truncates toward zero, `%` takes the sign of the dividend, + and `//` between floats divides (`7.5 // 2.0` is `3.75`). `cast(int)` rounds half to even + (`2.5` becomes `2`) and reads `'1.7'` as `2`. +* **Strings order by code point**, in comparisons, `min()` and `max()`, and `order_by` alike. +* **Names are matched without regard to case**, so a query is rejected when it is built if two + dimensions, or two fields of one struct, differ only by case. +* **A value an array cannot hold fails the query**: a failed `cast()`, a NaN or infinite result, a + null where the type is not optional, a repeated key, or a null dimension. Collecting an + `ArrayQuery` raises `QueryError`, from `fxtr.entity_defns.query`. In a workflow, the query node + fails, which fails the attempt as a failed step does; the workflow body cannot catch the error. + +## Keep query nodes stable + +**Put random draws and timestamps in steps, whose results are cached.** A query node is not +cached, so the runner computes it again when an attempt resumes or retries, and in a later job. +Computing it again gives the same answer, except for SQL that asks for varying values: +`random()`, `gen_random_uuid()`, `now()`, and the other clock functions. A sum of floating-point +values can also differ when the order of summation is not fixed, so put floating-point +aggregations whose exact value matters in a step too. + +When a resumed job, or a later job, gets a different answer, it cannot reuse what the earlier +run cached, and what happens depends on how the answer was used: + +* A step that receives the new answer is in cache conflict: the job stops with a cache conflict + at its address (see + [Resolving a job's conflicts](../concepts/managing-the-step-cache.mdx#resolving-a-jobs-conflicts)). +* A workflow that chose which steps to run from the answer can build a different graph. The job + then fails with `DeterminismError`, or, when no cached step is asked for a different call, + finishes with results computed from the new answer. + +To compute a varying value once, run the same query in a step with `await ... .collect()`. + +## Check a query's type + +`query.expect_type(...)` checks a query's type while it is built, as `ArrayHandle.expect_type` +does, so a wrong assumption about a query's shape fails where the query is written. For example, +`runs.select(value().field("score")).expect_type(("[model: str]", float))` raises +`ArrayTypeError`, because the result still has the `run` dimension. diff --git a/plugins/fxtr/skills/fxtr/references/reference/cli.mdx b/plugins/fxtr/skills/fxtr/references/reference/cli.mdx new file mode 100644 index 0000000..8cd69b5 --- /dev/null +++ b/plugins/fxtr/skills/fxtr/references/reference/cli.mdx @@ -0,0 +1,124 @@ +--- +title: "CLI reference" +description: "The fxtr command-line interface." +--- + +Run `fxtr` from inside a project with `uv run fxtr ...`. Every command works on the project +containing the working directory; `--project PATH` names another. `uv run fxtr COMMAND --help` +describes each command's options. + +| Global option | Effect | +| ----------------- | ----------------------------------- | +| `--project PATH` | A path inside the project to use. | +| `-q`, `--quiet` | Show only warnings from the runner. | +| `-v`, `--verbose` | Show the runner's debug output. | + +## Projects + +<ResponseField name="fxtr new DIRECTORY"> + Write an empty uv project that depends on this fxtr or a later one, with a Python package and an + empty renderer bundle. Add `--example` to include a runnable greeting experiment, its launcher, + greeting renderers, and a results overview. It runs `git init` (unless the directory is + already in a git or jj checkout) and prints the setup commands, starting with `uv sync`. + `--name` picks the project's name. +</ResponseField> + +<ResponseField name="fxtr init --database-url URL"> + Write `fxtr.local.toml`, which says where this machine reaches the project's database, and + create the project's tables. The URL is a SQLAlchemy `postgresql+asyncpg://...` URL. `--schema` + picks the database schema (by default `fxtr_` followed by the project's name), and `--force` + replaces an existing file. The `FXTR_DATABASE_URL` environment variable overrides the file's + URL. +</ResponseField> + +## Jobs + +<ResponseField name="fxtr run WORKFLOW"> + Run a workflow, by its registered name, as a job in this process until it stops. `--input + NAME=VALUE` passes an input as a JSON scalar, or as `@ARRAY` for a snapshot of a project array. + `--root` sets the root pathname. `--allow-staged` includes staged changes in the pinned code, + `--no-verify` skips checking the environment against `uv.lock`, and + `--overwrite-cache-conflicts` reruns steps in cache conflict and replaces the old results. + When the project's viewer is running, the job opens in a new tab of it; `--no-open` only + prints the link. +</ResponseField> + +<ResponseField name="fxtr resume JOB"> + Run a stopped job's unfinished work in this process until it stops. It reuses the results the + job already has and retries failed steps. Takes `--allow-staged`, `--no-verify`, + `--overwrite-cache-conflicts`, and `--no-open` like `fxtr run`. +</ResponseField> + +<ResponseField name="fxtr jobs list"> + List every job, newest first. +</ResponseField> + +<ResponseField name="fxtr jobs status JOB"> + Show where a job is, and why it stopped. A job's cache conflicts are shown grouped into + patterns, with the commands that resolve them. +</ResponseField> + +<ResponseField name="fxtr jobs cancel JOB"> + Ask the process running a job to stop it. +</ResponseField> + +<ResponseField name="fxtr jobs export JOB PATH"> + Write a stopped or succeeded job's record to a file: its facts, and every entity they reach, + as a gzipped CAR file that another project can import. +</ResponseField> + +<ResponseField name="fxtr jobs import PATH"> + Add the job in a record file written by `fxtr jobs export` to this project, with everything it + reaches. The commits its results were produced at must exist in this repository, and the job + must not be in the project already unless `--update` is given, which brings a job the project + has up to date from a later record of it. The step cache is left alone unless `--adopt-cache` + is given, which makes the cache offer the job's results as `fxtr cache adopt` does; with it, + `--overwrite-cache-conflicts` replaces entries that hold a different request's result. +</ResponseField> + +## Cache + +<ResponseField name="fxtr cache list [PATTERN]"> + List the step cache's entries, oldest first, with each entry's state, address, and step. + `PATTERN` limits the list to the entries at or under an address, where a mapped key's value + can be `*` (`'root/count_words[prompt:*]'`). `--conflicts-in JOB` lists the entries a job is + in cache conflict with, `--step` one step's entries, and `--state` entries in one state. +</ResponseField> + +<ResponseField name="fxtr cache clear [PATTERN]"> + Forget the entries at or under an address, so the next job to request them runs the steps + afresh. Jobs that already used a result keep it. `PATTERN` is as for `fxtr cache list`. + `--address` names an exact address, `--conflicts-in JOB` the completed entries a job is in cache + conflict with, and `--all` everything; `--step` and `--state` narrow any of these. +</ResponseField> + +<ResponseField name="fxtr cache adopt JOB [PATTERN]"> + Make the cache offer the results a job used, local or imported, so later jobs reuse them. + An address that holds a different request's result is left alone unless + `--overwrite-cache-conflicts` is given. `PATTERN` limits adoption as for `fxtr cache list`. +</ResponseField> + +## Viewer + +<ResponseField name="fxtr view"> + Serve the experiment viewer for the project's jobs and open it in a browser. It prefers the + project's last successful port, then tries 8000 through 8050. `--no-open` only prints the URL, + `--host` sets the bind address, and `--port` requires an exact port and fails if it is taken. + It builds the project's view bundle if the renderers changed + since the last build, then rebuilds it whenever they change while it runs, and open pages switch + to the new renderers without a reload. Failed builds keep the last valid bundle, or the + default renderers when none exists; the viewer shows the error and keeps watching for fixes. + When the project's viewer is already running, it opens that one and exits. +</ResponseField> + +<ResponseField name="fxtr views vendor"> + Copy the fxtr TypeScript packages the project's renderers build against into + `views/vendor/fxtr/`, replacing any earlier copy. Run it after upgrading fxtr, then + `pnpm install` in `views/`. Commit the copy with the project. +</ResponseField> + +<ResponseField name="fxtr views build"> + Build the project's renderer package into its view bundle, which `fxtr view` serves. A failed + build keeps the previous bundle. A running `fxtr view` rebuilds the bundle by itself; this is + for building without one. +</ResponseField> diff --git a/plugins/fxtr/skills/fxtr/references/reference/project-configuration.mdx b/plugins/fxtr/skills/fxtr/references/reference/project-configuration.mdx new file mode 100644 index 0000000..7cf2749 --- /dev/null +++ b/plugins/fxtr/skills/fxtr/references/reference/project-configuration.mdx @@ -0,0 +1,150 @@ +--- +title: "Project configuration" +description: "What fxtr new writes, and the settings in pyproject.toml and fxtr.local.toml." +--- + +A fxtr project is a uv project whose Python package defines steps and workflows. Two files beside +each other configure it: `pyproject.toml`, which is committed and shared by every checkout of the +project, and `fxtr.local.toml`, which says where this machine reaches the project's database and +is never committed. + +fxtr finds a project by walking up to the nearest `pyproject.toml`: from a launcher, starting at +the path it passes as its anchor (usually its own `__file__`), and from the `fxtr` command, at the +working directory (or `--project`). That file must have a `[tool.fxtr]` table; if it doesn't, fxtr +reports an error rather than looking further up. + +## What `fxtr new` writes + +`uvx fxtr new DIRECTORY` writes a project into a new or empty directory (from a project that +already depends on fxtr, `uv run fxtr new` does the same): + +```text +pyproject.toml the project, its dependency on fxtr, and [tool.fxtr] +README.md how to run it +.gitignore fxtr.local.toml, .env, .venv/, __pycache__/, .fxtr/ (where + `fxtr view` records its address), node_modules/, and dist/ +.env API keys for model providers +src/my_project/ + __init__.py +views/ the renderer package, with an empty bundle + vendor/fxtr/ the fxtr TypeScript packages the renderers build against +``` + +* The project is named after the directory, unless `--name` gives another name: `My_Project` is + the project `my-project`, whose code is the package `my_project`. +* It depends on the fxtr that ran `fxtr new`, or a later version. +* A new project has no experiment or launcher, and its `[tool.fxtr] modules` list is empty. + `--example` adds a greeting experiment in `src/<package>/experiment.py` (an entity type, a step, + and a workflow), registers that module, and adds `launch.py`, which runs it. It also + adds renderers for the greeting and a results overview of the workflow, + `views/src/overview.tsx`. +* `views/` has its own pnpm workspace boundary, lockfile, and installation, so installing its + dependencies never changes a workspace the project sits in. +* It runs `git init`, unless the directory is already in a git or jj checkout, and prints the + next steps. Run `uv sync` in the project to install its dependencies and write the virtual + environment and `uv.lock`. + +## `pyproject.toml` + +The file `fxtr new` writes, without its comments: + +```toml pyproject.toml +[project] +name = "my-project" +version = "0.1.0" +description = "A fxtr project" +requires-python = ">=3.13" +dependencies = ["fxtr[runner,debug-web]>=0.1.0"] + +[tool.uv] +package = true + +[build-system] +requires = ["hatchling"] +build-backend = "hatchling.build" + +[tool.hatch.build.targets.wheel] +packages = ["src/my_project"] + +[tool.fxtr] +modules = [] + +[tool.fxtr.views] +package = "views" +``` + +<ResponseField name="[tool.fxtr] modules" type="list of module names" default="[]"> + The modules whose import registers the project's steps and workflows. A process with no script + of its own imports them: a resumed job, `fxtr run`, and the viewer. Add each experiment module + you write to this list. A step or workflow that no listed module registers can still run in a + job launched by a script that imports it, but that job can't be resumed, `fxtr run` can't find + the workflow, and the viewer can't link its cards to the code. +</ResponseField> + +<ResponseField name="[tool.fxtr] stale_after" type="number of seconds" default="30"> + How long a job's lease may go without a heartbeat before fxtr presumes the process running it + is gone, and a resume may take the job over. Every process of a project must agree on it. +</ResponseField> + +<ResponseField name="[tool.fxtr.views] package" type="path"> + The directory of the project's renderer package, relative to the project root. Without this + table the project has no renderers of its own, and the viewer uses its defaults. +</ResponseField> + +The other settings are uv's and Python packaging's, chosen for fxtr: + +* `package = true` installs the project into its own virtual environment, so that the `fxtr` + command, which has no project directory on its import path, can import the modules in + `[tool.fxtr] modules`. Keep the definitions under `src/<package>/` and launcher scripts at the + project root. +* The `runner` extra is the Postgres runner, and `debug-web` is the viewer that `fxtr view` + serves. + +An experiment that calls language models also depends on +[behaviors](../../../behaviors/references/guides/language-models.mdx#set-up). A launch checks that the virtual +environment matches `uv.lock`, so commit `uv.lock` whenever the dependencies change (`uv add` and +`uv sync` update both). + +## `fxtr.local.toml` + +`uv run fxtr init --database-url URL` writes this file and creates the project's tables: + +```toml fxtr.local.toml +database_url = "postgresql+asyncpg://fxtr:fxtr@localhost:55433/fxtr" +schema = "fxtr_my_project" +``` + +<ResponseField name="database_url" type="string"> + A SQLAlchemy `postgresql+asyncpg://...` URL for the project's Postgres server. Required unless + the `FXTR_DATABASE_URL` environment variable is set, which overrides it. +</ResponseField> + +<ResponseField name="schema" type="string"> + The Postgres schema the project's tables live in: lowercase letters, digits, and `_`, not + starting with a digit, and at most 63 characters. By default it is `fxtr_` followed by the + project's name, lowercased, with every character other than a letter, digit, or `_` replaced by + `_`. `fxtr init --schema` writes another. +</ResponseField> + +`fxtr init --force` replaces an existing file. Everything the project stores lives in its schema: +jobs, project arrays, entities, and the step cache. Resuming or viewing a job therefore needs the +configuration that launched it. The project's client creates the tables when they are missing, but +it doesn't start the server or create the database. To keep trial runs apart from real results, +point the project at a scratch schema while experimenting (see +[Before you launch](../guides/running-and-viewing-jobs.mdx#before-you-launch)). + +fxtr reads both files strictly: a setting it doesn't know, or a value of the wrong type, is an +error, so a misspelled setting is never silently ignored. + +## `.env` + +`.env` holds the API keys of the model providers your experiments +call. `fxtr new` writes one with empty placeholders: + +```bash .env +ANTHROPIC_API_KEY= +OPENAI_API_KEY= +``` + +Fill in the keys you use and add any other variable your experiments read. +Keys in `.env` override variables set in the environment. diff --git a/plugins/fxtr/skills/fxtr/references/reference/renderers.mdx b/plugins/fxtr/skills/fxtr/references/reference/renderers.mdx new file mode 100644 index 0000000..dce7971 --- /dev/null +++ b/plugins/fxtr/skills/fxtr/references/reference/renderers.mdx @@ -0,0 +1,340 @@ +--- +title: "Renderers" +description: "The renderer API: slots and registration, loading entities and invocations, step panels, links, panel state, the kit, and building the renderer package." +--- + +A project's renderers are React components, written in TypeScript in the project's renderer +package, `views/`, and built into a view bundle that the viewer loads. +[Visualization and custom renderers](../concepts/visualization-and-custom-renderers.mdx) explains +how they fit into the viewer, with a small example, and +[Designing views](../guides/designing-views.mdx) covers how a view should look and behave. This page +is the reference for the API they use, `fxtr-view`, the contract with the viewer, and +`fxtr-view-kit`, the building blocks, and for the package that builds them. + +## Registering renderers + +`renderer(slot, subject, id, Component)` from `fxtr-view` registers a component for a slot and a +subject. The project's bundle, `views/src/bundle.ts`, exports its registrations as its default +`ViewBundle`: + +```ts views/src/bundle.ts +import { EntityPanel, InvocationPanel, renderer, type ViewBundle } from "fxtr-view"; +import { singleInvocation } from "fxtr-view-kit"; +import { DocumentPanel, Overview } from "./components"; + +const bundle: ViewBundle = { + renderers: [ + renderer(EntityPanel, "org.example.my_experiment.Document.v1", "my-project.DocumentPanel", DocumentPanel), + renderer(InvocationPanel, "my_experiment.run", "my-project.Overview", singleInvocation(Overview)), + ], +}; + +export default bundle; +``` + +| Slot | Subject | Props | What it draws | +| -------------------- | -------------------------------------- | -------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------- | +| `EntityPanel` | an entity type tag | `{ id, storage }` | The entity, opened in a pane. | +| `EntityInlineLink` | an entity type tag | `{ id }` | The entity where it is referenced, inside the viewer's link. It must be static text: no `EntityLink`, buttons, or other interactive elements. | +| `EntityHoverPreview` | an entity type tag | `{ id }` | A compact preview, shown when a link to the entity is hovered. | +| `InvocationPanel` | a step's or workflow's registered name | `{ invocations, storage }` | The invocations the viewer's selection picks out; for the root workflow, a view of the whole run. | + +The prop types are `EntityPanelProps`, `EntityViewProps` (the `id` alone), and +`InvocationPanelProps`. Give each renderer an ID prefixed with your project's name. Two renderers +for the same slot and subject with the same ID, including a clash with one of the viewer's own, +keep the whole bundle from loading. Several renderers with different IDs may target the same slot +and subject, and the viewer offers a switcher between them. + +Registrations belong in the bundle: a project never changes the viewer itself, its dependencies, +or its registry. + +## Loading data + +`unwrap(useEntity(id))` loads an entity's data, suspending until it arrives, so the component can +assume it is there and leave the loading state to the viewer. The data is the entity's stored +form: a map of its fields, in which references to other entities arrive as `EntityID`s. + +A view over many entities, such as a step's outputs, loads them with the kit's `LoadEach`, which +starts every load at once and hands its body the values loaded so far, by `idKey(id)`, with how +many are still pending, so the view fills in as its data arrives. `parse` turns each entity's +data into the value the body receives, and must be the same function on every render (defined +outside the component, or memoized). Never call `useEntity` in a loop. + +```tsx +<LoadEach ids={verdictIds} parse={parseVerdict}> + {(verdicts, pending) => <VerdictTable verdicts={verdicts} pending={pending} />} +</LoadEach> +``` + +An `InvocationPanel` is handed every invocation the viewer's selection picks out (its +`invocations` prop, each with its key in the run), and must show any number of them +([Step panels](#step-panels) says what each holds). The kit's +adapters turn a view of one loaded invocation record into such a panel: + +* `singleInvocation(View)` loads the record when there is one invocation and passes it to `View` + as `invocation`; for any other count it asks for one invocation to be selected. +* `eachInvocation(View)` shows several as collapsible sections, one per invocation, each loading + its own record when opened. + +A root workflow's panel is always given its one invocation. An invocation record has the +function's `name`, its `state` and `error`, its `inputs` by argument name, and its `output`, each +an array entity's ID to load in turn (`output` is null until there is one). The kit's +`invocationInput(invocation, name)` reads a named input, and `InvocationOutcome` shows a run that +has no output yet. + +## Step panels + +A step's `InvocationPanel` is its default view: the viewer draws it in the step's pane and as the +preview beside its card in the job's trace, so a reader sees it without asking. What it is given: + +* **Spawns, not records.** `invocations` holds one `Spawn` per invocation the selection picks out: + every invocation of the step, or fewer when the reader slices the level's dims (`slice dims`). + Handle any count. One spawn gets the entity it returned in full, through that entity's own + panel; several get a view of them all (a table, a grid, a chart) whose cells link to the entity + behind each. +* **A key of the mapped and replica dims only**, by name: `spawn.key.model`, `spawn.key.sample`. + `spawnKeyLabel(spawn)` from `fxtr-view-kit` joins them into a label. +* **An output array, null until the invocation has finished.** A step returning one entity + returns an array of one row with an empty key: `scalarEntityId(array)` reads the reference, and + `scalarValue(array)` reads a struct result. A step with output dims of its own + (`-> Annotated[Array[Hit], "[rank: int]"]`) returns those dims as the array's rows: read them + with `rowsOf(array)`, whose `key`s hold the output dims. So a spawn is keyed by the dims the + step is mapped over, and its output holds the dims the step introduces. +* **Inputs on the record.** `useInvocation(spawn.ref)` loads the spawn's `InvocationRecord`: its + `state` and `error`, and its `inputs`, array IDs by argument name + (`invocationInput(record, "scores")` from the kit). A dim the step takes whole (a core dim of an + argument, such as the `[doc]` a ranking step reduces) is in neither the key nor the output; + show it from the input. A hook is called a fixed number of times per component, so a view over + many spawns' records loads each in a component of its own, which fills in when its record + arrives. + +Load the outputs in two stages: the spawns' output arrays with one `LoadEach`, then the entities +they reference with another. A running step's spawns gain outputs as its invocations finish, and +`LoadEach` follows them, so the panel fills in while the step runs; say how many have no result +yet. A step mapped over `[model, prompt]` that returns a `Verdict` entity: + +```tsx views/src/judge.tsx +import { + EntityDataError, EntityLink, idKey, isArrayData, scalarEntityId, + type EntityData, type EntityID, type InvocationPanelProps, +} from "fxtr-view"; +import { InvocationHead, LoadEach, countLabel, formatNumber, layout, spawnKeyLabel } from "fxtr-view-kit"; +import { VerdictPanel, parseVerdict, type VerdictData } from "./verdict"; + +/** A step's output array. Module-level, so `LoadEach` gets a stable `parse`. */ +function asArray(data: EntityData) { + if (!isArrayData(data)) throw new EntityDataError(`expected an array, got ${data.$type}`); + return data; +} + +/** One verdict in full, or every verdict in a table that fills in as the invocations finish. */ +export function JudgePanel({ invocations }: InvocationPanelProps) { + const done = invocations.filter((spawn) => spawn.output !== null); + return ( + <div className={layout.panel}> + <InvocationHead kind="step" name="my_experiment.judge"> + <span className={layout.muted}> · {countLabel(done.length, "verdict")} of {invocations.length}</span> + </InvocationHead> + <LoadEach ids={done.map((spawn) => spawn.output!)} parse={asArray}> + {(arrays) => { + const rows = done.flatMap((spawn) => { + const array = arrays.get(idKey(spawn.output!)); + return array ? [{ spawn, id: scalarEntityId(array) as EntityID<VerdictData> }] : []; + }); + if (invocations.length === 1) { + return rows[0] ? <VerdictPanel id={rows[0].id} /> : <p className={layout.empty}>no verdict yet</p>; + } + return ( + <LoadEach ids={rows.map((row) => row.id)} parse={parseVerdict}> + {(verdicts) => ( + <table className={layout.table}> + <thead><tr><th>model · prompt</th><th className={layout.num}>score</th><th>verdict</th></tr></thead> + <tbody> + {rows.map(({ spawn, id }) => ( + <tr key={spawnKeyLabel(spawn)}> + <td>{spawnKeyLabel(spawn)}</td> + <td className={layout.num}>{formatNumber(verdicts.get(idKey(id))?.score)}</td> + <td><EntityLink id={id} /></td> + </tr> + ))} + </tbody> + </table> + )} + </LoadEach> + ); + }} + </LoadEach> + </div> + ); +} +``` + +Register it with `renderer(InvocationPanel, "my_experiment.judge", "my-project.JudgePanel", +JudgePanel)`. `singleInvocation` and `eachInvocation` suit a view of one invocation's record; a +panel over many spawns, like this one, reads them directly. + +## Links + +Anything that stands for an entity should lead to it. In text and tables, `EntityLink({ id })` is +the viewer's own link: a chip that opens the entity in a pane and shows its preview on hover. A +chart's mark can't be a chip, so it uses `useEntityLinks()`, whose `open(id)` does what pressing +the entity's link does, and whose `isOpen(id)` says whether a pane shows the entity. +[Designing views](../guides/designing-views.mdx#interaction) shows how marks use them. + +## Panel state + +A panel's `storage` prop is a JSON value the viewer keeps for the panel while it is open. Keep +what a reader chose in it (a tab, an order, a pick, a filter) with the kit's +`useStoredState(storage, key, initial, isValid?)`, which works like `useState` over one key of +it, so the choice survives the view remounting: + +```tsx +const [order, setOrder] = useStoredState(storage, "order", "guest"); +``` + +Stored values are trusted as the initial value's type unless you pass a type guard, `isValid`, to +reject an obsolete choice or state from another version of the view; a rejected value reads as +the initial one. The initial value and the guard are captured once for each storage and key. A view +wrapped with `singleInvocation` receives `storage` when its props declare it +(`{ invocation, storage }`). Keep hover and focus in `useState`. State a component keeps in +`useState` starts over when a rebuild replaces the renderers; state in `storage` stays. + +## Styles + +Style a renderer with CSS Modules (`*.module.css`) beside the components that use them, and with +the kit's `layout` classes, over the viewer's design tokens, so the view follows the viewer's +light and dark themes. A state or variant is a `data-*` attribute (`[data-state="failed"]`), not a +class name built from a value. [Designing views](../guides/designing-views.mdx) lists the tokens and +what each is for. + +## The kit + +`fxtr-view` is shared with the viewer, which hands its own copy to the bundle; `fxtr-view-kit` is +bundled with the project. Both are TypeScript source installed in the renderer package's +`node_modules/`: read `node_modules/fxtr-view-kit/src/` (or `fxtr-view/src/`) for a signature +this page leaves out, and run `pnpm typecheck` in `views/` (or `uv run fxtr views build`) to check +every use against them. + +From `fxtr-view`: + +| Name | What it is | +| -------------------------------------------------------- | ----------------------------------------------------------------------------------- | +| `renderer`, the slots, `ViewBundle` | registration ([above](#registering-renderers)) | +| `useEntity`, `unwrap`, `useInvocation`, `idKey` | loading an entity or an invocation record, and an ID's key | +| `isArrayData`, `scalarEntityId`, `scalarValue`, `rowsOf` | reading an array entity: a step's output ([Step panels](#step-panels)) | +| `EntityLink({ id })` | the viewer's link to an entity: a chip that opens it, with a hover preview | +| `useEntityLinks()` | `{ open(id), isOpen(id) }`: open an entity from a mark, and whether a pane shows it | + +From `fxtr-view-kit`: + +| Name | What it is | +| ---------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| `layout` | classes for elements you write: `panel`, `kindLine`, `title`, `subtitle`, `header`, `fieldsLine`, `prose`, `muted`, `empty`, `table` (the viewer's dense table), `num` (a number cell), `axis`, `grid`, `chartLabel`, `mark`, `cellMark`, `faded` | +| `Kind`, `InvocationHead`, `InvocationOutcome`, `FieldList` | a view's kind, an invocation's head, a run without output, label–value lines | +| `ChartFrame({ label, tooltip, fallbackWidth, children: (width) => ... })` | a chart's frame: measured width, and a tooltip `{ x, y, content }` in pixels from its top left | +| `useWidth(fallback)` | `[ref, width]` of any element, for charts in table cells | +| `pressable(onPress, label)` | props that make an element a keyboard-pressable, named button | +| `Legend({ items: [{ label, color, shape }] })` | series swatches (`shape`: `square`, `circle`, `line`) | +| `ScaleLegend({ scale, low, high, mid })` | a sequential or diverging scale's gradient between end labels | +| `seriesColor(i)`, `CHART_SERIES`, `CHART_OTHER` | categorical colors | +| `sequentialFill(t)`, `sequentialInk(t)`, `divergingFill(t)`, `divergingInk(t)` | scale colors, and text on them | +| `Stats`, `Stat({ label, value, note })` | a row of headline numbers | +| `Segmented({ label, options, value, onChange })`, `Select(...)` | controls; `options` are `{ value, label }` | +| `DetailBox({ title, onClose, children })` | the box showing a pick under a chart | +| `formatNumber(v, digits)`, `formatPercent(share, digits)`, `formatSigned(v, digits)`, `countLabel(n, noun)`, `MISSING` | numbers as the viewer writes them | +| `useStoredState(storage, key, initial, isValid?)` | `useState` kept in the panel's storage ([above](#panel-state)) | +| `singleInvocation`, `eachInvocation`, `invocationInput`, `LoadEach` | invocation adapters and loading ([above](#loading-data)) | +| `spawnKeyLabel(spawn)` | a spawn's key as a label ([Step panels](#step-panels)) | +| `MessageList`, `excerpt` | chat turns, and text cut to a length | + +The `fxtr-view-behaviors` package has views of [behaviors](../../../behaviors/references/index.mdx) conversation records, for a +project that stores them. + +## The renderer package + +`fxtr new` writes the renderer package, `views/`, with an empty bundle. `--example` fills the +bundle with a panel and an inline link for the example's greeting, and a results overview of its +workflow with a chart, `views/src/overview.tsx`. Install +its dependencies once, and commit the `views/pnpm-lock.yaml` it writes with your project: + +```bash +(cd views && pnpm install) +``` + +### Building and reloading + +`uv run fxtr view` builds the package into its view bundle when it starts if the bundle is missing +or stale, and again whenever you save a change to it while it runs. The open viewer then switches +to the new renderers in place, without a reload, keeping the job, level, and panes you have open, +and what each panel keeps in its `storage` ([Panel state](#panel-state)). + +To build without the viewer: + +```bash +uv run fxtr views build +``` + +This runs the package's build and, if the result checks out, replaces the bundle in `views/dist/`. +A failed build keeps the previous bundle and prints why it failed. Builds use the `node` on your +`PATH`, which must be Node 22 or later; fxtr refuses an older one and says how to fix it. + +The bundle is an ES module holding the renderers, an optional stylesheet, and a manifest naming +both and recording the version of the view API they were built against. The viewer loads it when +its page loads, checks that version against its own, and adds the bundle's renderers after its +defaults. React, jotai, and `fxtr-view` are not bundled in: the viewer hands its own copies to the +bundle, so renderers share the viewer's hook state and context. Everything else the renderers +import, the kit, a chart library, your own components, is bundled in. + +The fxtr packages the renderers import are not on npm. The project carries its own copy in +`views/vendor/fxtr/`, committed with the project. After upgrading fxtr, refresh the copy and +reinstall before building: + +```bash +uv run fxtr views vendor +(cd views && pnpm install) +uv run fxtr views build +``` + +### When the bundle doesn't load + +If a renderer build fails, including at startup, `fxtr view` keeps serving the last valid bundle, +or uses the default renderers when none exists. The side rail and terminal show the error and +the end of the build's output. Fix the source and save it to trigger another build; the viewer +updates without a restart. + +If a bundle fails to load in the browser, the viewer keeps the renderers it already loaded, or +uses its defaults when none were loaded, and the side rail says why. + +### Adding a package to a project + +To add a renderer package to a project without one, write these parts: + +* `[tool.fxtr.views] package = "views"` in `pyproject.toml`. +* `views/vendor/fxtr/`, the fxtr TypeScript packages: `uv run fxtr views vendor` copies them from + the installed fxtr. They are TypeScript source; inside `views/`, their own imports (React, + jotai, one another) resolve to `views/node_modules`. +* `views/package.json`, linking the fxtr packages from that copy + (`"fxtr-view": "link:./vendor/fxtr/packages/fxtr-view"`, and likewise `fxtr-view-kit`, + `fxtr-view-bundle` for the build, and `fxtr-view-behaviors` if you use it). Its build script is + `"build": "vite build --configLoader runner"`, whose flag lets the shared build preset's + TypeScript load; keep it when editing the script. React, jotai, TypeScript, Vite, and + `@types/node` are its development dependencies. `fxtr views vendor` names any registry package + a linked fxtr package needs that `package.json` lacks (`fxtr-view-shell-util` needs + `@ipld/dag-cbor` and `multiformats`). +* `views/pnpm-workspace.yaml`, with `packages: ["."]` and `onlyBuiltDependencies: [esbuild]`, which + gives the package its own installation and lockfile even inside another workspace, and allows + esbuild's install script. +* `views/vite.config.ts`, the build, which is always the shared preset's: + `export default defineViewBundleConfig({ entry: "src/bundle.ts" })`. +* `views/src/bundle.ts`, the entry, whose default export is a `ViewBundle`. +* `views/tsconfig.json`, extending `./vendor/fxtr/tsconfig.base.json`, which also declares CSS + Modules. +* `node_modules/` and `dist/` in the project's `.gitignore`. + +Then install its dependencies and commit the lockfile, as for a package `fxtr new` wrote. + +### Testing renderers + +If you test renderers with Vitest, set `resolve.dedupe` to +`["react", "react-dom", "jotai", "fxtr-view"]` in `vitest.config.ts`, so the linked fxtr packages +use the test environment's React and context instances.