Potato Documentation¶
Potato is a free, open-source, self-hosted annotation and agent-evaluation platform for NLP, agentic, and GenAI research. You configure tasks entirely in YAML — no coding — to annotate text, audio, video, images, documents, and AI agent traces, and to run a full agent-evaluation loop (programmatic evaluators, versioned datasets/experiments, automation, CI gating, LLM-as-judge calibration, and a multi-model arena) as a free alternative to LangSmith, LabelBox, and Braintrust.
New here? Start with the FAQ, the Glossary, or the Quick Start.
Looking for guides and tutorials?
This Read the Docs site is the complete, version-matched technical reference — every config option, the full HTTP API, and internals. For guided walkthroughs, use-cases, and higher-level docs, visit potatoannotator.com/docs.
Guides¶
Role-based guides that walk you through Potato for your specific use case:
- Getting Started Guide - First-time setup and your first annotation project
- Administrator Guide - Managing annotators, quality control, and monitoring
- Developer Guide - Extending Potato, API integration, and custom schemas
- Crowdsourcing Guide - Running tasks on Prolific and MTurk
- Agent Evaluation Guide - Evaluating AI agents, coding agents, and web agents
- AI-Assisted Annotation Guide - Using LLMs, active learning, and Solo Mode
Getting Started¶
- Quick Start - Get running in 5 minutes
- Installation & Usage - Detailed setup guide
- Reverse Proxy / URL Prefix - Run behind a path-prefix proxy (
/app1/) - Configuration Reference - Complete config options
- Comparison with Other Tools - How Potato compares to alternatives
For Coding Agents¶
Generated, machine-checkable specs built from the running code — prefer these over prose when writing a config or calling the API.
- Machine-Readable Specs - How to use all four, with editor setup and CI validation
- Config JSON Schema - Every valid
config.yamlkey, all 61 annotation types and 24 display types - OpenAPI 3.1 Spec - All 419 HTTP paths, with per-operation auth and config gating
- llms.txt - Curated documentation index (llms.txt standard)
- llms-full.txt - Every documentation page inlined into one file
Annotation Schemas¶
- Choosing the Right Annotation Type - Decision guide for selecting the best schema for your task
- Schema Gallery - All annotation types with examples
- Instance Display - Display images, video, audio, and text separately from annotation collection
- Conditional Logic - Show/hide questions based on prior answers
- Image Annotation - Bounding boxes, polygons, and landmarks
- Image Annotation Formats - Import COCO (polygons, RLE masks, crowd regions) with no preprocessing, and export back
- Deep Zoom & Tiling - Annotate images too large to send to a browser as one file, with masks painted at the source's full resolution and no GPU texture limit
- Point Cloud Annotation (3D) - Oriented 3D boxes on lidar: KITTI, PCD, PLY and LAS, converted server-side
- Calibration and 2D Verification - Project 3D boxes into the camera images, so a box is checkable rather than a guess
- Depth Maps - 16-bit PNG, NPY, PFM and EXR depth with a near/far window and a readout in metres
- Embodied Episode Annotation - Robot demonstrations: synchronized video, signal lanes, phases, outcomes and dense reward
- Robot Dataset Formats - LeRobot v2, RoboMimic/ALOHA HDF5, RLDS/TFDS, and Potato's own manifest
- Audio Annotation - Audio segmentation with waveform visualization
- Audio Dialogue - Podcast/interview turn annotation: speaker bubbles, per-turn audio playback, ratings, spans, and cross-turn linking
- Video Annotation - Frame-by-frame video labeling
- Tiered Annotation - ELAN-style hierarchical multi-tier annotation
- Transcript Formats - Whisper, subtitles, cloud ASR, TextGrid, and EAF: what Potato reads and how
- Working with Transcripts - From Whisper output or YouTube subtitles to a running annotation task
- Triage - Rapid accept/reject/skip data curation interface
- Entity Linking - Link spans to external knowledge bases (Wikidata, UMLS)
- Coreference Annotation - Group mentions of the same entity
- Conversation Tree Annotation - Annotate hierarchical conversation structures
- Format Support - PDF, Word, code, and spreadsheet annotation
- Span Linking - Relationship linking between text spans
- Soft Label - Probability distribution across labels via constrained sliders
- Confidence Annotation - Pair any annotation with an explicit confidence rating
- Constant Sum - Allocate a fixed budget of points across categories
- Semantic Differential - Bipolar adjective scales measuring connotative meaning
- Ranking - Drag-and-drop ordering of items by preference
- Range Slider - Dual-thumb slider for selecting an acceptable min-max range
- Hierarchical Multi-Label Selection - Select labels from an expandable tree taxonomy
- Visual Analog Scale (VAS) - Continuous analog scale for fine-grained magnitude estimation
- Extractive QA - SQuAD-style answer span highlighting
- Rubric Evaluation - Multi-criteria rubric grid for LLM evaluation
- Text Edit / Post-Edit - Inline text editing with diff tracking
- Error Span (MQM) - Error annotation with typed severity and quality scoring
- Card Sorting - Drag-and-drop grouping of items into categories
- Conjoint Analysis - Discrete choice between multi-attribute profiles
- Pairwise Comparison - Binary or scale-based A/B comparisons
- Multi-Dimensional Pairwise - Compare items on multiple axes simultaneously
- Best-Worst Scaling - Select best and worst from tuples
- Dialogue Annotation - Multi-turn conversation annotation
- Text Annotation - Free-text input and rationale annotation
- Event Annotation - N-ary event structures with triggers and arguments
- Trajectory Evaluation - Per-step error annotation for agent traces
Workflow & Quality¶
- Annotation Navigation - Navigation tools and status indicators
- Task Assignment - Assignment strategies and configuration
- Per-Cohort Schemas - Show different annotation schemes to different annotator cohorts
- Heterogeneous Coverage - Single-annotator default with a multi-annotator overlap sample, adaptive boost, per-annotator quotas, and full IAA reporting
- Diversity Ordering - Embedding-based clustering for diverse item presentation
- Multi-Document Event Annotation - Cross-document events with a 2D corpus map, cluster browser, KNN, and evidence-cited template slots
- Training Phase - Annotator training and qualification
- Quality Control - Attention checks and gold standards
- Boundary Lab - Counterfactual boundary probing: collect contrast sets during ordinary annotation, capture boundary rationales, and get paraphrase-invariance quality control
- Truth Serum - Surprisingly-popular scoring: gold-free item verdicts that beat majority vote, plus annotator calibration
- Paper Mode - One command generates a cut-paste LaTeX dataset report: description, distributions, annotator table, IAA, limitations
- Think-Aloud Mode - Voice rationales with fully-local STT and rule-based spoken-label commitment; verbatim reasoning streams, no LLM
- Pocket Mode - Mobile-first card-stack annotation PWA: thumb-zone labels, swipe navigation, offline queue with auto-sync
- Psychometrics - Labels with error bars: live IRT (ability + difficulty, no gold, no LLM), information-gain adaptive routing, codebook-bug detection, and pre-study power analysis
- Multiplayer Rooms - Live group annotation: blind-vote norming sessions with a real-time agreement meter and conformity logging, adjudication huddles, and expert shadowing
- Adjudication - Multi-annotator disagreement resolution
- MACE - Multi-Annotator Competence Estimation via variational inference
- Iterative BWS - Adaptive Best-Worst Scaling for fine-grained ordinal rankings
- Category Assignment - Category-based item assignment
- Surveyflow - Pre/post annotation surveys
- Annotation Filtering - Filter data based on prior annotations
- Survey Instruments - 55 pre-built validated psychological instruments
- QDA Mode - Qualitative data analysis workspace: composes codebook + memos + cases + search with single-coder defaults
- Memos - Universal annotator notes (instance/span-anchored, private/shared)
- Search - Universal FTS5 search; admin search + guarded annotator search-and-claim
- Codebook - Universal mutable code set (nested, opt-in per scheme, on-the-fly add)
- Cases - Group instances into units of analysis; QDA auto-detect; crosstab integration
Agent Evaluation¶
- Coding Agent Annotation - Evaluate agentic coding systems (Claude Code, SWE-Agent, Aider) with diff rendering, PRM annotation, and code review
- CoT Process Reward (LLM pre-label + verify) - Segment a long chain-of-thought into steps, have an LLM pre-label each step's reward, and have a human verify — fast PRM data collection
- VLM Grounding & Pointing - Bind referring expressions to regions or points, with an explicit not-present answer, and localize what a caption says that is not there
- World-Model & Rollout Evaluation - Frame-locked video panels; mark the frame where a generated rollout stops making sense, and measure whether annotators agree on when
- Agent Traces - Evaluate AI agent traces and trajectories
- Turn-Level Annotation - Bind any rating/tagging/comment schema per-turn with declarative filters (by speaker, agent, step type, tool)
- Multi-Agent Discussion - Annotate agent-to-agent discussions/debates with agent identity, addressees, reply threading, and consensus tracking
- Agent Task Recipes - Ready-to-run configs: debate judging, plan review, negotiation, safety escalation, context-use annotation
- Session-Level Scoring - Group traces by session_id/thread_id and score whole sessions on a dedicated queue page
- Sub-Agent Run Tree - Interactive run-hierarchy sidebar for orchestrator traces; bind per-turn schemes to specific sub-agent runs
- Reviewer Routing + Kanban - Route instances to reviewers with first-match rules; track review states on a kanban board with adjudication handoff
- Three-Pane Trace Eval - Reasoning | function calls | final answer side-by-side, for continuous evaluation
- Trajectory Correction - Edit traces into SFT/DPO training data
- Datasets & Experiments - Versioned eval datasets + experiment runs that score outputs over time
- Programmatic Evaluators - Trajectory match, tool-use, LLM-judge & heuristic evaluators (Flask-free library)
- Automation Rules - filter→sample→action rules that route incoming items to queues/datasets/evaluators (production→eval loop)
- CI Evaluation - pytest plugin to run evals in your suite and gate the build on score-threshold regressions
- Semantic Curation - embedding search + dynamic slices to find traces by similarity and curate them into datasets
- LLM-Judge ↔ Human Alignment - Measure & calibrate an LLM judge against human gold (Cohen's κ)
- Signal-Based Triage Queue - Prioritize the queue by a quality signal (errors / low score first)
- Hotkey Review Mode - Keyboard-driven review queue with auto-advance on completion
- Live Agent Interaction - Observe and interact with a live AI agent in real time
- Model Arena - Compare N models side by side on one prompt; pick the best, build a win-rate leaderboard (provider-agnostic)
- Web Agent Annotation - Review and create web agent browsing traces
Solo Mode¶
- Solo Mode - Human-LLM collaborative annotation workflow
- Solo Mode Advanced Features - Edge case rules, labeling functions, confidence routing
- Solo Mode Developer Guide - Architecture and extension points
AI & Intelligence¶
- AI Support - AI-powered label suggestions
- Using HuggingFace Models - Point AI hints, solo mode, and judge calibration at any HF model
- Judge Calibration - Auto-label with LLM judges + blind human calibration (accuracy, IAA, ECE)
- Active Learning - ML-based prioritization
- Active Learning Strategies - Query strategies reference (BADGE, BALD, hybrid, cold-start)
- ICL Labeling - In-context learning for labeling
- Visual AI Support - YOLO and vision LLM support for image/video annotation
- Chat Support - LLM-powered sidebar for annotator assistance
- Option Highlighting - AI-assisted highlighting of likely annotation options
- Embedding Visualization - UMAP-based instance similarity dashboard
Authentication & User Management¶
- Users & Collaboration - User registration, access control, and collaboration
- Roles & Permissions (RBAC) - Role-based access control: role→permission mapping, per-user and SSO role assignment
- Password Management - Password security, reset flows, database backend, and shared credentials
- Passwordless Login - Authentication without passwords
- SSO & OAuth Authentication - Google, GitHub, and institutional SSO login
Crowdsourcing¶
- Crowdsourcing Guide - Prolific and MTurk integration
- MTurk Integration - Detailed Amazon MTurk setup guide
Administration¶
- Scaling & Large Datasets - How Potato handles big datasets, indexing, memory, and bulk exports
- Admin Dashboard - Monitoring and management
- Annotator Progress Dashboard - Opt-in, read-only progress view for annotators
- Behavioral Tracking - User behavior analytics
- Keystroke Logging - Content-blind typing dynamics on free-text fields
- Writing-Process Detection - Tell composed text from transcribed or LLM-pasted text
- Keystroke Logging Ethics - IRB, consent, and participant rights
- Annotation History - Tracking annotation changes
Data & Output¶
- Data Format - Input and output data formats
- Export Formats - Export to COCO, YOLO, CoNLL, and more
- Publishing Datasets - One-click publish to HuggingFace, Zenodo (DOI), or a documented archive with an auto-generated dataset card
- HuggingFace Hub Export - Push annotations to HuggingFace Hub
- HuggingFace Datasets Integration - Load annotations as DatasetDict or DataFrame
- Remote Data Sources - Load data from S3, Google Drive, Dropbox, URLs, and databases
- Data Directory - Load data from a directory with optional live watching
UI & Customization¶
- UI Configuration - Interface customization
- Layout Customization - Custom CSS layouts and styling
- Form Layout - Grid layout, column spanning, styling, and alignment
- Multilingual - Localization and RTL support
Integrations¶
- Webhooks - Outgoing webhook notifications for annotation events
- HuggingFace Spaces - Deploy Potato on HuggingFace Spaces
- LangChain Integration - Send LangChain agent traces to Potato
- Tracing SDK (
potato_trace) - Instrument any agent with@traceableto capture runs into Potato (OpenTelemetry interop)
Tools & Utilities¶
- Preview CLI - Preview configs without running server
- Migration CLI - Upgrade v1 configs to v2
- Debugging Guide - Debug flags and troubleshooting
- Simulator - Annotation simulation tool
- API Reference - REST API endpoints documentation
Productivity Features¶
- Productivity - Tooltips, shortcuts, and highlights
Release Notes¶
- v2.7.1 - Transcripts In, Without the Reformatting (21 input formats, sidecar files,
potato transcriptsCLI) - v2.7.0 - Seven New Ways to Annotate (Psychometrics, Multiplayer Rooms, Boundary Lab, Truth Serum, Think-Aloud, Paper Mode, Pocket Mode)
- v2.6.1 - Agentic Evaluation Suite (evaluators, datasets/experiments, automation, CI gating, tracing SDK, curation, arena)
- v2.6.0 - QDA Mode, LLM-as-Judge Calibration & Trajectory Editing
- v2.4.4 - Span Annotation Fixes & UX Improvements
- v2.4.3 - Coding Agent Annotation, Localization & Stability
- v2.4.1 - Bug Fixes
- v2.4.0 - Agent Evaluation, AI-Assisted Annotation & Enterprise Integration
- v2.3.0 - Solo Mode, Agent Workflows & Security Hardening
- v2.2.0 - Comprehensive Annotation & Export Platform
- v2.1.0 - Adjudication & Multi-Modal Annotation
- v2.0.0 - Backend Refactor
Contributing¶
- Contributing Guide - How to contribute to Potato
Quick Links¶
| Task | Documentation |
|---|---|
| Set up a basic annotation task | Quick Start |
| Choose an annotation type | Schema Gallery |
| Display images/video with radio buttons | Instance Display |
| Show/hide questions based on answers | Conditional Logic |
| Annotate PDFs, Word docs, or code | Format Support |
| Set up SSO/OAuth login | SSO Authentication |
| Reset a user's password | Password Management |
| Use a database for user storage | Password Management |
| Configure for MTurk | MTurk Integration |
| Configure for Prolific | Crowdsourcing Guide |
| Monitor annotation progress | Admin Dashboard |
| Add AI suggestions | AI Support |
| Set up quality control | Quality Control |
| Present items diversely | Diversity Ordering |
| Debug configuration issues | Debugging Guide |
| Create custom visual layouts | Layout Customization |
| Rapidly filter/triage data | Triage |
| Link entities to knowledge bases | Entity Linking |
| Annotate coreference chains | Coreference Annotation |
| Annotate conversation trees | Conversation Tree Annotation |
| Navigate efficiently through items | Annotation Navigation |
| Evaluate AI agent traces | Agent Traces |
| See reasoning, tool calls & answer side-by-side | Three-Pane Trace Eval |
| Edit agent traces into SFT/DPO training data | Trajectory Correction |
| Build versioned eval sets & track scores over time | Datasets & Experiments |
| Score agent outputs programmatically | Programmatic Evaluators |
| Capture agent runs from your own code | Tracing SDK |
| Auto-route incoming traces to queues/datasets/evals | Automation Rules |
| Gate CI on eval-score regressions | CI Evaluation |
| Find traces by similarity / curate slices | Semantic Curation |
| Compare models side by side on a prompt | Model Arena |
| Align/calibrate an LLM judge to human labels | Judge Alignment |
| Auto-label with LLM judges + calibrate blind | Judge Calibration |
| Use Solo Mode for collaborative annotation | Solo Mode |
| Export annotations to Parquet | Export Formats |
| Export to COCO/YOLO/CoNLL | Export Formats |
| Import COCO annotations (RLE masks included) | Image Annotation Formats |
| Push annotations to HuggingFace Hub | HuggingFace Export |
| Deploy on HuggingFace Spaces | HuggingFace Spaces |
Run behind a /app1/ reverse proxy |
Reverse Proxy / URL Prefix |
| Set up webhook notifications | Webhooks |
| Use LLM chat assistant for annotators | Chat Support |
| Evaluate coding agents | Coding Agent Annotation |
| Review web agent traces | Web Agent Annotation |
| Load data from S3/GDrive/Dropbox | Remote Data Sources |
| Use pre-built survey instruments | Survey Instruments |
| Set up multilingual interface | Multilingual |
| Resolve annotator disagreements | Adjudication |
| Browse the REST API | API Reference |
Example Projects¶
Ready-to-use example configurations are available in the examples/ directory:
# Run a simple radio button example
python potato/flask_server.py start examples/classification/single-choice/config.yaml -p 8000
# Run a custom layout example (content moderation, dialogue QA, medical review)
python potato/flask_server.py start examples/custom-layouts/content-moderation/config.yaml -p 8000
See the examples directory for more examples, including:
classification/- Classification annotation examples (radio, checkbox, likert, etc.)span/- Span annotation examples (NER, linking, coreference, etc.)image/- Image annotation examplesvideo/- Video annotation examplesaudio/- Audio annotation examplesadvanced/- Advanced features (conditional logic, quality control, etc.)agent-traces/- Agent trace evaluation examples (RAG, GUI agents, comparisons)custom-layouts/- Custom layout examples
Citation¶
If you use Potato in your research, please cite the Potato 2.0 paper (ACL 2026 System Demonstrations):
@inproceedings{jurgens-etal-2026-potato,
title = "Potato 2.0: A Comprehensive Annotation Platform with {AI}-in-the-Loop Support",
author = "Jurgens, David and Chen, Michael and Iyer, Lina",
booktitle = "Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 3: System Demonstrations)",
month = jul,
year = "2026",
address = "San Diego, California, United States",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2026.acl-demo.37/",
pages = "374--386",
}
The original Potato release is described in the Potato 1.0 paper (EMNLP 2022 System Demonstrations, Pei et al., 2022). See the README for both BibTeX entries.