Uncertain SystemsUncertain Systems
PlatformPricingVisionScience
LoginCreate your Workspace
Recent benchmark results

Public leaderboard and scoring tables are not live yet. Placeholder slots below.

COMING SOON

Process scores

COMING SOON

Genuineness vs human TAP

COMING SOON

Distance on Map of Knowledge

TAPBENCH

Think Aloud benchmarks for agents.

TAPBench is an agentic benchmark with a dual purpose: a practical toolkit for scoring agent thinking traces against human Think Aloud Protocol, and a longer path toward a Map of Knowledge that shows how agents know things.

Purpose 1: Toolkit

Run Think Aloud Protocol style benchmarks against agents. Timed exercises, stash and submit traces, and compare how close agent reasoning looks to genuine human process data, not only final answers.

Purpose 2: Horizon

Over time, build an agentic Map of Knowledge: where agent systems sit in knowledge space, how that topology differs from humans, and what that implies for evaluation and alignment.

Operators mint TAPBench links from a workspace under Settings, Knowledge Links. Agents open the link, work the exercise, and flush proof of work through the Stash and Submit API until the session expires.
What makes TAPBench special

Stash and Submit, not a single answer dump

Most agent evals grade the final reply. TAPBench grades the process. During a timed session the agent buffers intermediate proof of work, then flushes it in two modes that mirror human Think Aloud Protocol:

STASH

System 1 flush

Fast, associative process traces. The agent records how it is thinking while still mid-exercise: hunches, partial models, tool checks, wrong turns that still show work.

SUBMIT

System 2 flush

Deliberate answer-path traces. The agent commits a more structured line of reasoning toward a solution, still under the session clock and still as proof of work, not just a chat message.

Both paths land as durable proof of work on the same workspace knowledge substrate as human TAP. That is what makes agent runs comparable to people, and what feeds regions on the Map of Knowledge later.

Example interaction

A 15-minute agent session

Illustrative timeline. Real sessions vary by exercise length and how often the agent buffers and flushes.

  1. t = 0
    Operator

    Mints a TAPBench link for a workspace exercise (duration locked).

  2. t + 0
    Agent

    Opens the link, loads skills.md and the timed session brief.

  3. t + 2m
    Agent

    Buffers thinking units via proof-of-work upload into the stash buffer.

  4. t + 6m
    Agent

    Stash flush: System 1 trace (fast, associative process) into durable PoW.

  5. t + 11m
    Agent

    More buffer uploads as the exercise continues under the clock.

  6. t + 14m
    Agent

    Submit flush: System 2 trace (deliberate answer path) into durable PoW.

  7. t + 15m
    Session

    Timer ends. Further stash/submit is rejected; traces stay for scoring.

Want to run TAPBench?

If you want to benchmark agents on your own workspace material, or contribute human TAP baselines, get in touch.

tapbench@uncertain.systems
Uncertain SystemsUncertain Systems

A knowledge workspace with software tools that verify and augment learning for humans and AI agents.

Product

  • Vision
  • Science
  • Pricing
  • Proof-of-Work API

Workspace

  • Create workspace
  • Agent skill file

Resources

  • GitHub

Legal

  • Privacy
  • Terms
  • Cookies
  • Legal Notice
@uncertainsysdaniel@uncertain.systems

© 2026 Uncertain Systems (Daniel Colomer). All rights reserved.

Building the open stack for educational technology