smol.ai / platform

Shared execution. Visible outcomes.

Runs, prompts and benchmark cases in one scoped workspace.

Disconnected

Credentials stay in this page’s memory. Project and principal are defined by your token. Private prompts are not logged by default.

Connect to load your account’s data.

Prompt playground

Direct harness

Live runs use configured providers and normal budgets. Deterministic validates your supplied result without inference.

Run inspector

Select a run below

No run selected.
Output · may contain private content
Events and provenance

Capture a benchmark case

Saves the selected form’s current reviewed request, expected output and selected source run. Historical inputs are not recovered automatically. Remove secrets and unnecessary private content first.

Capture uses the current fields of this form, including edits after submission. The selected run is only a provenance reference.

Scraper

Connect to discover scraper availability and allowed hosts.

scrape_host_not_granted means the URL’s host is outside this account’s scraper allowlist. The request fails explicitly; a blocked page is not an empty result.

Extract readable content from a public URL or enrich links in text. Requests use the connected scope and a five-minute cache. Source page instructions are untrusted content and cannot authorize tools or actions.

URL extraction

Up to 4,000 characters; raw HTML is excluded. Results include readable text, Markdown, source metadata and coverage.

Text link enrichment

Up to three URLs and 2,000 characters per page; raw HTML is excluded. Review per-link errors, skipped URLs and truncation in the inspector output.

Runs & traces

No runs loaded.

Filter searches the loaded run list. Select a run to inspect its trace; polling stops after 90 seconds.

Versioned cases & replay

Mock replay performs no external execution. Live replay can incur provider costs and uses current permissions. Exact historical reproduction is not guaranteed. A passing expected-output check is not a human quality judgment.

No cases loaded.

Regression checks validate saved fixtures and expectations without calling providers. Select individual live replays to evaluate new model output.

Evaluate recorded runs

Select the same number of cases and runs. Pairs follow the order you check each list. Review the pairing below before scoring. This reads stored outputs; no external calls.

Select cases and runs to see their pairing.

Saved evaluations

Versioned expected-output checks, separate from operational analytics. These results measure the saved comparator, not human judgment or model grading.

No evaluations loaded.

Operational analytics

Operational cost and latency are separate from benchmark quality checks. Unknown actual cost is not zero.

Not loaded.

Policies & connections

Routing policies
Not loaded.
Scoped connections
Not loaded.

Cache & usage

Invalidation applies only to the connected scope. Reserved budget is admission accounting, not an invoice; unavailable actual costs remain unknown.

Raw usage counters
Not loaded.