Quick Start
This guide walks you through running your first Scrivai PES.
Prerequisites
- Python >= 3.11
- Scrivai installed (
pip install scrivai) - Claude Agent SDK CLI available (
claudecommand) - API key configured (via
.envfile or environment variable)
pip install scrivai
# Create .env with your API credentials
echo 'ANTHROPIC_BASE_URL=https://your-gateway.example.com' >> .env
echo 'ANTHROPIC_AUTH_TOKEN=your-key-here' >> .env
Minimal AuditorPES Example
The following script audits a document against a set of checkpoints.
import asyncio
from pathlib import Path
from pydantic import BaseModel
from scrivai import (
AuditorPES, ModelConfig, WorkspaceSpec,
build_workspace_manager, load_pes_config,
)
# 1. Define output schema
class AuditOutput(BaseModel):
findings: list[dict]
summary: dict
# 2. Set up workspace
ws_mgr = build_workspace_manager()
spec = WorkspaceSpec(
run_id="quickstart-audit",
project_root=Path("."), # must contain skills/ and agents/ dirs
data_inputs={"document.md": Path("my_document.md")},
force=True,
)
ws = ws_mgr.create(spec)
# 3. Place checkpoints in workspace
import json
checkpoints = [
{"id": "CP001", "description": "Document must have a title"},
{"id": "CP002", "description": "All figures must have captions"},
]
(ws.data_dir / "checkpoints.json").write_text(
json.dumps(checkpoints, ensure_ascii=False)
)
# 4. Load config, create PES, and run
config = load_pes_config(Path("scrivai/agents/auditor.yaml"))
model = ModelConfig(model="claude-sonnet-4-20250514")
pes = AuditorPES(
config=config,
model=model,
workspace=ws,
runtime_context={"output_schema": AuditOutput},
)
async def main():
run = await pes.run("Audit data/document.md against all checkpoints")
print(f"Status: {run.status}") # "completed"
print(f"Findings: {run.final_output}") # {"findings": [...], "summary": {...}}
asyncio.run(main())
The Three-Phase Lifecycle
Every PES run goes through exactly three phases:
┌──────────────────────────────────────────────────────────┐
│ PES Run │
│ │
│ ┌──────────┐ ┌──────────────┐ ┌───────────────┐ │
│ │ PLAN │───▶│ EXECUTE │───▶│ SUMMARIZE │ │
│ │ │ │ │ │ │ │
│ │ Read │ │ Per-item │ │ Merge all │ │
│ │ inputs, │ │ processing │ │ findings into │ │
│ │ produce │ │ with tool │ │ output.json │ │
│ │ plan.json│ │ use │ │ │ │
│ └──────────┘ └──────────────┘ └───────────────┘ │
│ │
│ phase_results phase_results phase_results │
│ ["plan"] ["execute"] ["summarize"] │
└──────────────────────────────────────────────────────────┘
- Plan: The Agent reads inputs and generates
plan.json+plan.md. - Execute: The Agent follows the plan, producing
findings/<id>.jsonper item. - Summarize: The Agent merges findings into
output.json(validated by the framework).
Each phase is recorded as a PhaseResult containing all turns and the final text.
Document Conversion (OCR)
Scrivai provides to_markdown() for quick one-off conversion and MarkdownConverter for production pipelines with multi-backend dispatch, concurrent processing, and fault tolerance.
Quick Conversion
from scrivai import to_markdown
# Zero-config: reads SCRIVAI_GLM_API_KEY / SCRIVAI_OCR_BASE_URL from env
md = to_markdown("report.pdf")
print(md)
MarkdownConverter (Recommended)
For full control over backends, concurrency, and error handling:
from scrivai import MarkdownConverter
# Single backend
with MarkdownConverter(monkey_endpoints=["http://gpu6:7866"]) as conv:
md = conv.convert("report.pdf")
# Multi-backend with overflow dispatch
with MarkdownConverter(
monkey_endpoints=["http://gpu6:7866", "http://gpu7:7867"],
glm_api_keys=["key1"],
) as conv:
md = conv.convert("report.pdf")
# GPU mode (local MinerU)
with MarkdownConverter(mode="gpu") as conv:
md = conv.convert("report.pdf")
Batch Processing
with MarkdownConverter(monkey_endpoints=["http://gpu6:7866"]) as conv:
results = conv.convert_batch(["a.pdf", "b.pdf", "c.pdf"])
# results: {"a.pdf": "...", "b.pdf": "...", "c.pdf": "..."}
Supported Formats
| Input | Pipeline |
|---|---|
.pdf |
OCR backend (MonkeyOCR / GLM-OCR / MinerU) |
.docx |
markitdown (local, no OCR) |
.doc |
LibreOffice → .docx → markitdown |
CLI
# API mode (default)
scrivai-cli io convert --input report.pdf --output report.md
# GPU mode (local MinerU)
scrivai-cli io convert --input report.pdf --output report.md --mode gpu
For the full architecture guide — dispatch strategies, fault tolerance, chunking, and configuration reference — see Concepts: IO.
Next Steps
- Concepts: PES Engine — file contracts, prompt templates, custom PES
- Concepts: Workspace — run isolation and
extra_env - Concepts: Trajectory — record and replay runs
- API Reference: PES — full class documentation