Reasoner Design & Decisions
How the reasoner (S2) actually thinks — the connective narrative behind the
individual decisions. Where reasoner.md is the reference (the
contract, the cadence, the env vars) and each decision is recorded in isolation in
the private management decision log, this page is organized by the logic
problem the reasoner has to solve. For each problem: what it is, how we solve
it, and why we chose that.
It deliberately does not restate mechanism detail that lives in
reasoner.md or the openral_reasoner_ros README —
it links to them.
One sentence: the reasoner is an event-driven LLM supervisor that closes
context → LLM → one typed tool callat ~0.2 Hz; it decides what to do next and never drives motors — it proposes, the C++ safety kernel disposes.
A recurring design through-line runs under every section below: wrap an unreliable LLM in deterministic, bounded scaffolding. The LLM decides what is true and what to try; typed state machines, calibrated signals, and hard caps decide what is recorded and when to stop. Keep that in mind — it explains most of the choices.
1. The core loop — one typed tool call per slow tick
Problem. A robot supervisor that calls an LLM on a fixed fast timer burns tokens doing nothing, and a free-form LLM that emits prose or multiple simultaneous actions can't be safely dispatched or replayed.
Solution. The reasoner is event-driven with a slow heartbeat:
- Heartbeat at
tick_hz = 0.2(one tick / 5 s). A heartbeat tick that has seen no new event short-circuits insideReasonerCorewithsuppressed_reason="heartbeat_idle"— no LLM call, no span. - Event preemption is the real trigger, gated by a hard
min_interval_s = 100 msso nothing can thrash the LLM. Four tiers:
| Tier | Source topic | Preempts on |
|---|---|---|
| A — safety | /openral/failure/safety |
severity ≥ WARN |
| B — execution | /openral/failure/{hal,sensor,rskill,wam} |
severity ≥ FAIL |
| C — progress | /openral/failure/critic |
severity ≥ FAIL |
| D — operator/world | /openral/prompt, /openral/perception/* |
new prompt forces; perception is informational |
- Each tick the LLM emits exactly one variant of the
ReasonerToolCalldiscriminated union (ExecuteRskill,LifecycleTransition,ReloadGstPipeline,EmitPrompt, plus the read-only query/memory tools below) — Pydantic-validated structured output, never free-form JSON.
Execution model (#21). The blocking LLM round-trip (select_tool, and the
VLM gate's describe_image — each on its own dedicated single worker, so a
safety tick never queues behind an adjudication call) runs off the rclpy
executor. A tick is phased: ReasonerCore.prepare_tick (gates +
context render, executor thread) → run_prepared_llm (worker) →
finish_tick (bookkeeping + dispatch, marshaled back via a guard condition).
One worker = one outstanding LLM call; ticks requested mid-flight coalesce and
replay (the single-flight trampoline). This is what makes "event-driven
preemption" real: while the LLM thinks for 10–60 s, goal results, patience
timers, and Tier-A preemptions still run on time (pre-#21, a 10 s patience
timer was observed firing at 76 s because the timer callback queued behind the
LLM call). Events that arrive mid-flight were not in the model's context, so
finish_tick marks seen/drains only the prepare-time snapshot — a mid-flight
operator prompt survives for the next tick.
Authority boundary. The reasoner holds no actuation authority: it never
publishes ActionChunk. Only rskill_runner_node does, and every action passes
the C++ safety kernel before it
reaches a motor. Python proposes; C++ disposes.
Why. Event-driven cuts idle LLM calls ~85% vs a fast timer; the heartbeat is the deadlock insurance ("task not progressing"). One typed call per tick is what makes a run dispatchable, traceable, and replayable from the trace alone.
2. Knowing what's in front of it — perception without 3D
Problem. To act on "the cup" the reasoner must know a cup is visible and be
able to refer to it across ticks. The 3D scene graph (scene_objects) needs
depth to lift object poses — but many deploy cameras are RGB-only (LIBERO),
so scene_objects comes up empty and the reasoner has nothing to ground against.
This is exactly what made a real run loop forever on a collective goal.
Solution. Two complementary perception surfaces, both read-only, neither requiring depth:
- Camera-space
in_viewenumeration. The continuous detector stamps a stable per-objectdet_id(via a 2D-IoUDetectionTracker2D) and the context renders a line the LLM can refer to:in_view[top]: #0 milk @px(412,233), #1 ketchup @px(388,251), …. Pixel centers, not 3D poses — kept in a separate line fromscene_objects[map]: …@(x,y,z)so coordinate spaces never blur. Identity exists with or without depth. - A sticky
locatedline. The continuous detector's fixed ~230-class vocab mislabels the goal objects (a basket read as "box", ketchup as "bottle"). So every successful open-vocablocate_in_viewhit is folded byContextRenderer.note_located()into a persistentlocated[<cam>]line (latest-wins,_LOCATED_CAP=12). The prompt tells the LLMlocatedis authoritative over the noisyin_view.
On-demand localization. The
locate_in_view tool asks a live detector "is X in camera Y right now?" via the
/openral/perception/<detector>/locate_in_view service. Detectors run
node-per-detector so a continuous detector and one or more on-demand locators
coexist, and LocateInViewTool gets a detector selector (fast
omdet-turbo-locator for simple "find X", locateanything-3b for referring
expressions). A recall_object miss auto-escalates to a live locate_in_view
before handoff — policy in the node, not dependent on the LLM picking the tool.
Two hard-won usability fixes (as of the 2026-06-29 amendment): (1)
omdet-turbo-locatoris a multi-label detector — query it with concrete object nouns / a comma-list ("cup, bowl, basket"), never a collective phrase ("the objects on the table"), which it matches as one nonexistent class. (2) The launch setsprimary_camera=det_cameraso frames cache under the real camera name ("top"); otherwise everylocate_in_view(camera="top")missed against a frame stored under"default". Both surfaced as the samefound=Falseloop.
Active search (proposed). When an object isn't in memory at all, recall_object /
resolve_place plus a bounded SearchBudget (max places and wall-clock,
no hidden default) drive a look → navigate → re-query loop, opening occluding
containers first, terminating in human-handoff.
Why. The reasoner doesn't need 3D poses to decide — it needs labels + a
stable id to refer. Camera-space enumeration is cheap (it already subscribes to
the detector) and depth-free. The sticky located line closes the real gap where
the cheap continuous detector mislabels but the open-vocab locator confirms.
3. Turning a goal into actionable subtasks — grounded decomposition
Problem. A collective operator goal ("put all the objects on the table into the basket") has to become an ordered list of single-object subtasks. Live testing showed that prose asking the LLM for "one specific object per subtask" isn't enough — weak models emit vague "batch" subtasks even with the grounded object list in context.
Solution.
- A structural contract, not a prompt.
DecomposeMissionTool.subtasksislist[GroundedSubtask], whereGroundedSubtask(object_ref, text)carries a Pydantic@model_validatorthat rejects a collectiveobject_ref/text(sharedis_collective_targetpredicate) and requirestextto nameobject_ref. The type makes a vague subtask un-representable on the wire. - A sequential task queue.
MissionState(tasks, current)holdsTaskStates with a strict lifecyclepending → active → verifying → {done|abandoned}(at most one active). The operator goal seeds a single task (MissionState.from_prompt); the LLM decomposes viaDecomposeMissionTool, which populates the same queue. A blocked task can be subdivided in place (subdivide_active), bounded byDEFAULT_MAX_SUBDIVIDE_DEPTH = 2.
Why. "Types are the contract" (CLAUDE.md §1.3) — a structural invariant the model can't violate beats prose it can ignore. The queue gives the reasoner a deterministic record of where it is, so progress doesn't depend on the LLM re-deriving the plan every tick. The decomposition itself still needs a capable model — see §8 (Choosing the brain) below.
4. Knowing when a subtask is actually done — reward-gated, VLM-adjudicated completion
Problem. A VLA emits action chunks but no notion of success. The skill
runner's result.success only means "the policy ran to its deadline without
crashing," not "the task was accomplished." Gating completion on a clock, or on a
single hardcoded 0.8 threshold, can't tell getting closer from stuck and
misclassifies a physically-successful result scored 0.78.
Solution — a layered signal stack:
- A reward model running parallel to the VLA.
kind: rewardrSkills (default Robometer-4B, NF4, ~3.6 GB) score the shared camera stream every ~1–2 s and exposeprogress_now/success_now/ trends through the read-onlyquery_task_progresstool. ItsRewardContractmanifest block declares the calibration (success_threshold,frame_window_s, …). Advisory only. - A stall watchdog that fires a stream, not a poll.
Every reward model publishes self-describing
CriticScoreon/openral/critic/score; acritic_id-keyedCriticWatchdogGroupwatches for a stall and publishes a Tier-CFailureTriggeron/openral/failure/critic— so a plateau preempts a tick instead of silently running to timeout. - A reward-watcher wake.
The instant the reward signal hits success, plateau, or the
patience ceiling, the in-flight VLA is cancelled and a normal reasoner
tick wakes with the reward trajectory injected.
patience_s(anExecuteRskillToolfield, default from the contract) replaces the LLM-guesseddeadline_sas the execution backstop. - A three-tier verdict.
evaluate_task_verdictreplaces the hardcoded threshold: - auto-pass (
score ≥ success_threshold) →complete_active, no VLM call; - vlm_check (
check_floor ≤ score < success_threshold) → adjudicate the current frame withdescribe_image("is<task>complete? yes/no"); - ladder (
score < check_floor) → no VLM, straight to replanning.
Amendment 2026-06-29 — gate on the progress head, not success: progress separates genuine success (0.80–0.86) from failure (~0.74), while the success head is compressed/noisy (0.56–0.79). Both heads are rendered to the LLM in a
## REWARDcontext block (set_reward_state). The reward also scores the whole attempt (start→now), not an 8 s trailing window (frame_window_sraised8.0 → 40.0). A per-taskTaskLocateBudgetabandons afterDEFAULT_MAX_TASK_LOCATE_ATTEMPTS = 3locate cycles that never dispatch a VLA.
Why. The reward signal is already calibrated and continuous — a clock and a single threshold throw that away. The authority stack (system fallback < reward contract default < LLM per-task override) scales to future per-task SARMs with no re-architecting. Degradation is honest: with no reward oracle and no VLM, the ambiguous middle is never claimed as success — it runs to the patience ceiling and hands off.
5. Always running a VLA with its reward model — pairing + VRAM fit
Problem. §4 only works if the reward model is actually co-resident with the VLA. Pairing used to be an implicit deploy flag decoupled from which VLA the reasoner picks at runtime, and nothing guaranteed both fit on the GPU before loading — you'd discover the mismatch as a mid-run CUDA OOM.
Solution. A VLA
manifest names its reward model (reward_rskill_name, allowed only for
kind == "vla"; None = deployment default) and declares per-dtype VRAM
(min_vram_gb, read by active_min_vram_gb()). A pure helper
assert_vla_reward_fits(vla, reward, gpu_total_gb, margin_gb=0.5) raises
ROSConfigError (undeclared) or ROSGPUMemoryError (won't fit). It runs at two
points: the reasoner's _refuse_unfittable_vla drops a non-fitting VLA from the
palette so it's never dispatched, and (defense-in-depth) the runner re-checks
before from_pretrained. The deploy CLI adds a pre-launch preflight
(_preflight_reward_vram_fit, torch-free nvidia-smi probe) that hard-exits only
when no capability-matched VLA fits — deploy sim / deploy run only, not
benchmark / sim run. (Eviction of other peers — detectors before the VLA —
is handled by a complementary decision.)
Why. A VLA without a reward model is blind to its own success, so it should never run alone. Sizes are knowable from the manifests; an oversized pair should fail before launch with an actionable message, not as an opaque OOM mid-grasp.
6. When things go wrong — the bounded replanning ladder
Problem. A failed or stalled step must escalate through progressively more disruptive recovery, and must be guaranteed to terminate (no infinite retry storm).
Solution. A fixed ladder:
retry → param-tweak → substitute-skill → goal-replan → human-handoff. The
shipped gate is ReasonerCore's per-kind retry cap (retry_cap_per_kind,
default 3) — consecutive same-kind selections beyond the cap are suppressed
(suppressed_reason="retry_cap"), and the streak resets on a material context
shift. At the mission layer, TaskState.attempts bounds a task's total tries;
exhaustion calls abandon_active, emits an honest "could not complete task K"
with the MissionState snapshot, and advances. The terminal rung is always
human-handoff.
Why. Bounded everything (CLAUDE.md §1.4) — every recovery path has an explicit
cap and no hidden default, so a wrong reward reading or a stuck skill costs a
finite number of tries, then surfaces honestly. (The substitute-skill and
goal-replan rungs are partially realized; the retry cap + attempts cap + handoff
ship today — see reasoner.md §Bounded replanning.)
7. Learning and reusing knowledge — playbooks + self-maintained memory
Problem. The reasoner has strong mechanism but thin content: decision procedures lived as bespoke Python, it didn't learn across episodes, and it never saw its own body or its execution outcomes.
Solution (proposed/phased):
- Playbooks (
kind: "playbook",role: "s2") — Markdown SOPs the LLM reads and interprets, never executes. TheirPlaybookContract(trigger,composes_tools,done_predicate,max_steps, fallback) is selection metadata; thePLAYBOOK.mdbody is injected into the## PLAYBOOKSprompt block. Six launch playbooks encode the recurring procedures (find-object,decompose-mission,verify-outcome,preflight-reach,stage-for-manipulation,clarify-ambiguity). - Self-maintained
MEMORY.md— semantic/narrative memory (preferences, corrections, lessons, durable home facts), complementary to the geometric scene graph: the scene graph answers "where is the mug?",MEMORY.mdanswers "how does this household like things done?". The LLM edits it only through typedMemoryWriteTool/MemorySearchToolops (add/update/supersede/delete) — a traced event, never a free-form rewrite — and a periodic consolidation pass keeps it bounded. - Context grounding — a
## ROBOTself-model block (reach hull, FOV, gripper, control modes, derived fromRobotCapabilitiesat configure), a## EXECUTIONblock (one NL line per skill outcome, success and failure), and the## REWARDblock from §4. These close the loop so the next tick reasons on reality.
Why. Playbooks make decision procedures authored content (versioned,
shipped, discoverable) instead of code. Memory is split by kind (geometric vs
semantic) so each lands in its correct consumer. Every write is a discrete traced
call — the planner can't silently corrupt the file, and a run stays replayable
from the trace + the MEMORY.md snapshot. None of this adds actuation authority.
8. Choosing the brain — LLM selection & the deploy default
Problem. The reasoner's hardest job (grounded decomposition of a collective
goal, §3) is genuinely reasoning-heavy. Weak/cheap models follow the one-tool-
per-tick contract fine but over-locate and never call decompose_mission;
the library must stay provider-agnostic (no cloud lock-in).
Solution. Selection is model-first (ADR-0088). OPENRAL_REASONER_MODEL
names a curated ReasonerModel; the registry resolves dialect, endpoint, auth,
hosting, and local-compute requirements. Endpoint location is orthogonal via
OPENRAL_REASONER_ENDPOINT, so Ollama/vLLM/cloud placement is not encoded in a
misleading provider enum. The library factory has no default and refuses
to guess. The deploy-sim launch defaults to the curated gpt-5.5 entry, with
OPENRAL_REASONER_MAX_TOKENS=16384 so a reasoning model does not reserve its
full window and get 402'd on a metered key. The default needs an API key and
fails loudly without it. Raw uncurated models require an explicit endpoint +
dialect and emit a warning; the old provider-first env remains a one-release shim.
Why. In live deploy testing GPT-5.5 was the only model that reliably
decomposed the collective goal (glm-5.2 over-located and never decomposed; Opus
4.8 worked but needed nudges; the OpenRouter :free tier emitted the placeholder
skill id). Simpler single-object goals run fine on the cheaper baselines in the
README.
Design through-lines
The same principles recur across every section — if you internalize these, the individual decisions follow:
- Deterministic scaffolding around an unreliable LLM. Typed
MissionState,GroundedSubtask, calibrated reward gates, and hard caps do the bookkeeping the model can't be trusted to. The LLM decides what is true and what to try; the scaffolding decides what is recorded and when to stop. - Everything bounded, no hidden defaults (CLAUDE.md §1.4). Heartbeat, min-interval, retry cap, attempts cap, subdivide depth, search budget, patience ceiling, locate budget — all explicit.
- Perception is advisory; the kernel disposes (CLAUDE.md §1.1). Detectors, reward models, critics, scene graph, and memory are all read-only inputs to a decision; none commands a motor. The C++ safety kernel is the only authority.
- Honest degradation. Missing a reward oracle or a VLM never produces a claimed-uncertain success — the attempt runs to the bound and hands off.
- Types are the contract (CLAUDE.md §1.3). A vague subtask or an unpaired reward model is made un-representable on the wire, not merely discouraged.
See also
reasoner.md— the reasoner reference (cadence, tool contract, palette gating, provider table, observability, how to run it).openral_reasoner_rosREADME — ROS wrapper contract, provider presets, baseline LLM configs.- Architecture overview · repo state map.