Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
12 changes: 12 additions & 0 deletions .agents/plugins/marketplace.json
Original file line number Diff line number Diff line change
Expand Up @@ -15,6 +15,18 @@
"authentication": "ON_INSTALL"
},
"category": "Productivity"
},
{
"name": "fxtr",
"source": {
"source": "local",
"path": "./plugins/fxtr"
},
"policy": {
"installation": "AVAILABLE",
"authentication": "ON_INSTALL"
},
"category": "Productivity"
}
]
}
11 changes: 10 additions & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,6 +7,14 @@ A collection of plugins for Codex.
| Plugin | Description |
|--------|-------------|
| [docent](./plugins/docent) | Docent AI analysis tools for Codex |
| [fxtr](./plugins/fxtr) | Skills for writing, running, and viewing fxtr experiments, and for calling models from them with behaviors |

The fxtr plugin includes two skills, each with a copy of the documentation it links from
[docs.transluce.ai](https://docs.transluce.ai):
- **fxtr** - Setting up, writing, running, and viewing fxtr experiments
- **behaviors** - Calling language models from fxtr experiments with the behaviors library

The fxtr plugin is generated from the fxtr repository by `pnpm sync:plugin`; don't edit it by hand.

## Installation

Expand All @@ -16,8 +24,9 @@ Add this marketplace to Codex:
codex plugin marketplace add TransluceAI/codex-plugins
```

Then install the plugin:
Then install a plugin:

```shell
codex plugin add docent@transluce-plugins
codex plugin add fxtr@transluce-plugins
```
23 changes: 23 additions & 0 deletions plugins/fxtr/.codex-plugin/plugin.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,23 @@
{
"name": "fxtr",
"version": "0.1.0",
"description": "Write, run, and view fxtr experiments, and call models from them with behaviors.",
"author": {
"name": "Transluce",
"url": "https://transluce.org"
},
"homepage": "https://docs.transluce.ai/fxtr",
"license": "Apache-2.0",
"keywords": ["fxtr", "behaviors", "experiments", "skills"],
"skills": "./skills/",
"interface": {
"displayName": "fxtr",
"shortDescription": "Write, run, and view fxtr experiments.",
"longDescription": "Skills for writing, running, and viewing fxtr experiments, and for calling models from them with behaviors.",
"developerName": "Transluce",
"category": "Productivity",
"capabilities": ["Read", "Write"],
"websiteURL": "https://docs.transluce.ai/fxtr",
"defaultPrompt": ["Use fxtr to set up a new experiment project.", "Use fxtr to sweep this prompt over several models and judge the replies."]
}
}
38 changes: 38 additions & 0 deletions plugins/fxtr/skills/behaviors/SKILL.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,38 @@
---
name: behaviors
description: Call language models from fxtr experiments with the behaviors Python library, covering one-shot requests and judge steps, multi-turn chat rollouts, scripted or custom user and tool policies, retries and failures, streaming, checkpoints, and storing conversations. Use whenever a fxtr experiment calls a model or holds a conversation, or when extending behaviors' model adapters or policies.
---

# behaviors

behaviors is a Python library for calling language models from fxtr experiments: provider
adapters, conversation sessions, context policies that play users and tools, retries and failure
recording, and conversation records stored as fxtr entities. Use these pieces when they fit the
experiment, rather than building another chat loop or transcript format.

The pages below are its documentation. Read the ones a task needs before writing code.

## Where to read

| Task | Read |
| ------------------------------------------------------------------------------------------------------------------------------ | ---------------------------------------------------------------------------------- |
| Add behaviors to a project; hold a conversation, judge outputs, and summarize verdicts in fxtr steps; test without credentials | [Calling language models](references/guides/language-models.mdx) |
| Pick a provider, an adapter, and an exact model ID | [Choosing models](references/guides/choosing-models.mdx) |
| Script users, offer tools, or write a custom conversation flow | [Policies and tools](references/guides/policies-and-tools.mdx) |
| Rollout statuses and `max_turns`, failures, retries, streaming | [Rollout outcomes and retries](references/guides/rollout-outcomes.mdx) |
| Messages and content, stored records, transcripts, conversations in the viewer, checkpoints | [Conversation records and checkpoints](references/guides/conversation-records.mdx) |
| Run Docent readings, or convert records for Docent | [Docent readings](references/guides/docent-readings.mdx) |
| Choose a layer; imports, adapters, sampling settings, requests and responses | [API overview](references/reference/api.mdx) |
| Steps, workflows, arrays, launching jobs, and the step cache | the adjacent [fxtr skill](../fxtr/SKILL.md) |

## Best practices

* **Judge with `gpt-6-luna` by default.** For an LLM judge, use `gpt-6-luna` through the
`OpenAIResponsesAPI` adapter, unless the user asks for another model.
* **Keep the model the user names.** Verify its exact ID from the provider's listing, as Choosing
models shows, rather than substituting another. Without credentials to check, take candidate
IDs from the provider's published catalog, and tell the user that account access has not been
verified.
* **Shape model experiments as the guide does**: one step per conversation, with every setting
that affects the request in a configuration entity; fxtr replicas for repeated samples; and
separate steps for judging and for aggregating verdicts.
11 changes: 11 additions & 0 deletions plugins/fxtr/skills/behaviors/provenance.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,11 @@
{
"skill": "behaviors",
"documentation": "https://docs.transluce.ai/behaviors",
"packages": {
"fxtr": "0.0.1a0",
"behaviors": "0.0.1a0"
},
"revision": "216737e51a92f5a0a83f7629af73d5a733fceb19",
"modified": false,
"content_sha256": "de0ca958d570464776b006b595d0f26b7be280a6c5baf88c46629d2982757eee"
}
19 changes: 19 additions & 0 deletions plugins/fxtr/skills/behaviors/references/INDEX.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,19 @@
# The behaviors documentation

The pages of https://docs.transluce.ai/behaviors, copied with this skill: `provenance.json`,
beside its `SKILL.md`, says from which version. Each is listed with what it covers,
in the order of the site's navigation.

## Get started

- [What is the behaviors library?](index.mdx): Primitives for evaluating AI model behavior.

## Other pages

- [Choosing models](guides/choosing-models.mdx): Find the exact model IDs a provider offers your account, pick the adapter for its endpoint, and check a model before a sweep.
- [Conversation records and checkpoints](guides/conversation-records.mdx): What a rollout records, how conversations are stored as fxtr entities and flattened into transcripts, and how to checkpoint a session.
- [Docent readings](guides/docent-readings.mdx): Convert conversations to Docent's data models, and run a Docent reading against a transcript or agent run with a model you choose.
- [Calling language models](guides/language-models.mdx): Use the behaviors library to hold conversations and judge outputs in fxtr steps.
- [Policies and tools](guides/policies-and-tools.mdx): Decide what the model sees each turn: scripted users, tool policies, their combinations, and custom conversation flows.
- [Rollout outcomes and retries](guides/rollout-outcomes.mdx): How a rollout ends, how model call failures are retried, recorded, or raised, and how to watch a rollout as it runs.
- [API overview](reference/api.mdx): The layers of behaviors, where each is imported from, the provider adapters, and the request, response, and sampling types.
110 changes: 110 additions & 0 deletions plugins/fxtr/skills/behaviors/references/guides/choosing-models.mdx
Original file line number Diff line number Diff line change
@@ -0,0 +1,110 @@
---
title: "Choosing models"
description: "Find the exact model IDs a provider offers your account, pick the adapter for its endpoint, and check a model before a sweep."
---

behaviors takes a provider's model ID as given: it has no catalog of models, and it passes the
string to the provider's endpoint without checking it against a list or translating it. Choose
the provider and endpoint first, then get the exact model ID from that provider's listing API.
Use the returned `id`, not a display name; the same model can have different IDs at different
providers. When you've been asked for a particular model, keep it, and verify its ID rather than
substituting another.

The examples below use `curl` and `jq`, with API keys already set in the named environment
variables. They only list model metadata. Use the same credentials and endpoint as the
experiment, since a listing shows what that account can use. An authentication or permission
error does not mean that no models exist.

## OpenAI

[List models](https://developers.openai.com/api/reference/resources/models/methods/list)
with `GET https://api.openai.com/v1/models`:

```bash
curl --fail-with-body --silent --show-error https://api.openai.com/v1/models \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-o /tmp/openai-models.json &&
jq -r '.data[].id' /tmp/openai-models.json | sort
```

Use a returned ID with `OpenAIResponsesAPI(model_id)` for models supporting the Responses API, or
`OpenAICompatibleAPI(model_id)` for Chat Completions. The listing includes models for other
tasks too; check a model's supported endpoints and features in the
[model catalog](https://developers.openai.com/api/docs/models) before choosing an adapter.

## Anthropic

[List models](https://platform.claude.com/docs/en/api/models/list)
with `GET https://api.anthropic.com/v1/models`:

```bash
curl --fail-with-body --silent --show-error 'https://api.anthropic.com/v1/models?limit=1000' \
-H "x-api-key: $ANTHROPIC_API_KEY" \
-H 'anthropic-version: 2023-06-01' \
-o /tmp/anthropic-models.json &&
jq -r '.data[] | [.id, .display_name] | @tsv' /tmp/anthropic-models.json &&
jq '{has_more, last_id}' /tmp/anthropic-models.json
```

This is one page, even with `limit=1000`. While `has_more` is true, request the next page with
`after_id` set to the returned `last_id`, and collect its models before advancing again.
Pass a returned `id` to `AnthropicMessagesAPI(model_id)`; `display_name` is only for presentation.

## OpenRouter

OpenRouter serves models from multiple vendors through an OpenAI-compatible endpoint. Its
[general catalog API](https://openrouter.ai/docs/api/api-reference/models/list-all-models-and-their-properties)
is `GET https://openrouter.ai/api/v1/models`; the
[account-filtered listing](https://openrouter.ai/docs/api/api-reference/models/list-models-filtered-by-user-provider-preferences-privacy-settings-and-guardrails)
also applies your provider preferences, privacy settings, and guardrails:

```bash
curl --fail-with-body --silent --show-error https://openrouter.ai/api/v1/models/user \
-H "Authorization: Bearer $OPENROUTER_API_KEY" \
-o /tmp/openrouter-models.json &&
jq -r '.data[] | [.id, .name] | @tsv' /tmp/openrouter-models.json
```

For the general catalog, change `/models/user` to `/models`. The
[model browser](https://openrouter.ai/models) is useful for interactive search. Inspect
`architecture`, `context_length`, and `supported_parameters` in the JSON for candidate models.
Use the full returned `id`, including its vendor prefix and any suffix, with:

```python
import os

from behaviors.implementations.models.openai_compatible import OpenAICompatibleAPI

api = OpenAICompatibleAPI(
model_id, # Exact id selected from the OpenRouter listing.
base_url="https://openrouter.ai/api/v1",
api_key=os.environ["OPENROUTER_API_KEY"],
)
```

The OpenAI SDK does not read `OPENROUTER_API_KEY` by itself, so pass it explicitly as above.
Close the adapter with `await api.aclose()` when finished, as in the
[conversation step](language-models.mdx#a-conversation-as-a-step).

## Search and validate a candidate

Search a saved response by ID or display name, changing the file and search text as needed:

```bash
jq -r --arg query 'claude' '
.data[]
| select(([.id, (.name // .display_name // "")] | join(" ") | ascii_downcase)
| contains($query | ascii_downcase))
| .id
' /tmp/openrouter-models.json
```

For paginated responses, search every collected page. A catalog entry does not establish that
the model supports every feature the experiment needs. Check its support for the endpoint,
streaming, tools, and reasoning, then try the chosen adapter and sampling settings with a small
request before launching a sweep. Without credentials, the providers' published catalogs and
documentation give candidate IDs, but not whether your account can use them.

Keep the chosen provider, model ID, and request settings in the experiment's configuration
inputs, so that each run records what was selected (see
[A conversation as a step](language-models.mdx#a-conversation-as-a-step)).
Original file line number Diff line number Diff line change
@@ -0,0 +1,81 @@
---
title: "Conversation records and checkpoints"
description: "What a rollout records, how conversations are stored as fxtr entities and flattened into transcripts, and how to checkpoint a session."
---

A rollout returns a `ChatRollout`: its context turns, its model turns, how it ended, and the
initialization its model session started from. This page covers what those records hold, how they
are stored and shown, and how they differ from a session's checkpoint.

## Messages and content

Messages are `SystemMessage`, `UserMessage`, `AssistantMessage`, and `ToolMessage`, each holding a
tuple of typed content blocks. An assistant's content may include text, tool calls, reasoning,
redacted reasoning, or summarized reasoning, so don't assume every block has `.text`. For the
answer's text alone, keep the `ContentText` blocks of the turn's `AssistantMessage`s, as the
[judge step's](language-models.mdx#a-judge-step) `answer_text` does, and keep the
structured rollout for later inspection.

A model turn made by an API-backed session records the provider's response in
`ModelTurn.metadata["response"]`, as JSON, including its stop reason and usage. Metadata is defined
by whatever produced the turn, so validate it before scoring with it.

## Storing records

`ChatRollout`, `Transcript`, `TranscriptGroup`, `AgentRun`, and the concrete session
initializations, such as `SetSystemMessage`, are fxtr entities, with type IDs of the form
`org.transluce.behaviors.<Class>.v1`. A step that returns a `ChatRollout` has it stored like any
other entity (see [Entities and custom types](../../../fxtr/references/concepts/entities-and-custom-types.mdx)), and
`await storage.store(entity)` and `await storage.load(ref)` work on them too. To load a bare ID,
pass the concrete type, as in `as_type=ChatRollout`, not an abstract base such as
`SessionInitialization`. A tree of transcripts and groups is stored as one entity, so there is no
need to store each child separately.

fxtr checks stored records strictly: give fields their declared types (`temperature=0.0`, not
`0`), and keep metadata to JSON values with string keys.

## Transcripts

`ChatRolloutTranscriptProjection`, from `behaviors.rollouts.chat`, flattens a rollout into a plain
`Transcript`: `await ChatRolloutTranscriptProjection().to_transcript(rollout)` gives one message
list, starting with the system message of a recorded `SetSystemMessage` initialization and
followed by the turns in order, including a final context turn the model never answered.
`to_agent_run` wraps that transcript in an `AgentRun`. The projection drops each turn's metadata,
so judge from the native rollout when you need usage, stop reasons, or other metadata.

## Showing conversations in the viewer

The `fxtr-view-behaviors` package draws rollouts, transcripts, transcript groups, agent runs, and
system-message initializations as conversations. To use it, spread its `behaviorsRenderers` into
your project's bundle:

```ts views/src/bundle.ts
import type { ViewBundle } from "fxtr-view";
import { behaviorsRenderers } from "fxtr-view-behaviors";

const bundle: ViewBundle = {
renderers: [...behaviorsRenderers],
};

export default bundle;
```

A new project's renderer package doesn't depend on it yet: add
`"fxtr-view-behaviors": "link:./vendor/fxtr/packages/fxtr-view-behaviors"` to the dependencies in
`views/package.json`, like the other fxtr packages, and run `pnpm install` in `views/` (see
[The renderer package](../../../fxtr/references/reference/renderers.mdx#the-renderer-package)). The package also exports
`ChatRolloutView`, for a rollout inside a view of your own, and `ComparedRollouts`.

## Checkpointing sessions

A serializable session can save its state and be restored from it. `await session.get_state()`
returns a JSON object, and `await factory.restore_session(state)` starts a session in that state.
The sessions of `chat_session_model_from_api` and of `ScriptedUserSimulator` support this, and the
`serializable_context_policy_from_*` combinators build checkpointable policies when every component
they combine is checkpointable. Generator-based policies can't be checkpointed, because a
suspended generator can't be saved.

A checkpoint is not a stored record: it is a session's working state, apart from the transcript
entities and from fxtr's step cache. fxtr doesn't resume a partly finished rollout from a
checkpoint: a step that failed partway runs again from the start when its job resumes. Close the
sessions you start yourself with `aclose()`.
Loading
Loading