The Economy as an RL Environment: Notes on Competence and Consequence
The economy is not one game. It is many connected practices. Each practice has its own rules for evidence, obligations, and payoffs. To call an agent competent means it can move through these different rule systems without flattening them into a single number or a single trick. As Wittgenstein put it, "meaning is use." In our context, the meaning of a model's output depends on how a professional can actually use it.
Reinforcement learning needs more than a reward signal. In institutional life, success is not just "win." It is "produce the thing as professionals expect it." A trade ticket, a regulatory filing, a valuation memo are not abstract scores. They are deliverables judged against standards that already exist. Aristotle reminds us, "We are what we repeatedly do." Competence therefore shows up in repeated, acceptable deliverables, not in one clever move.
Rules in these settings are not just limits. They are the living form of the work. To draft a valuation memo is to enter a practice with expectations about sources, logic, disclosure, and risk. An RL environment becomes economic when tasks reflect the real structure of work: ambiguity that does not disappear, deadlines that force choices, steps that depend on earlier steps, objectives that can conflict, and the duty to ground claims in shared evidence. Kant's insight helps here: "Thoughts without content are empty." A model's reasoning needs the content of the domain to count as judgment.
Consequence is the teacher. A system that never meets the friction of procedure will learn shortcuts, not judgment. It must face what qualifies as proof, which risks matter, and what cannot be said without citation. ZORA encodes these grammars into evaluation environments where agents are judged for competence. The question is simple. Can the model produce deliverables whose structure, grounding, and usefulness meet the standards of finance and law?
The conclusion is practical. If intelligence is to be useful, it must live inside real practices. We design environments where doing the work as professionals do it becomes the path to improvement. Reward becomes recognition. A good memo is accepted for review. A solid advisory can be acted upon. As Popper wrote, "All life is problem solving." ZORA turns professional problems into learning signals so that models learn to solve the kind of problems that actually matter.
Deliverable-Level Evaluation: A Grammar of Utility
To distinguish an answer from a deliverable is to distinguish speech from action. An answer satisfies curiosity; a deliverable satisfies obligation-the obligation to enable a decision within a practice governed by rules, tools, and consequences. In this sense, utility is not a decorative property appended after performance; it is the grammar within which performance has meaning. What counts as a "good" artifact is inseparable from the form of life that uses it.
Conventional benchmarks assay isolated capacities-classification accuracy, chain‑of‑thought fluency, code synthesis. But a valuation memo or legal advisory is not a single capacity; it is a composition. Evidence must be gathered, normalized, and cited; arguments must be structured; risks must be surfaced, ranked, and explained; recommendations must connect analysis to an actionable path. The criterion of correctness therefore becomes plural. Accuracy without grounding misleads. Structure without consequence trivializes. Eloquence without risk awareness deceives. Utility is the fit between artifact and practice: whether the output can be accepted, amended, or deployed by a professional under constraints of time, regulation, fiduciary duty, and scrutiny.
Rubrics must speak the language of the domain: what counts as sufficient citation, which models are defensible, where sensitivity should be shown, how uncertainty is narrated without theatricality. Inter‑rater calibration does more than reduce variance; it constructs a stable grammar in which judgments become teachable signals for models. Disagreement, properly mapped, reveals the edges of practice-the places where expertise is most instructive.
ZORA's deliverable‑level evaluation makes models confront the obligations of professional life. Scores cease to be a talent show and become a readiness index: can this artifact be used? The final remark is plain: utility is grammar. To evaluate intelligence is to ask whether it can produce artifacts that withstand the practitioner's test. When models learn this grammar, they no longer imitate answers; they compose decisions.
Alignment Signals as Forms of Life: On Teaching Models Judgment
"Alignment" is often described as steering models toward preferred behavior, but preference is too weak a word for professional domains. A form of life-the practice of law, the craft of finance-sustains itself through obligations: evidential standards, procedural disciplines, risk doctrines, and a habit of justification. To teach judgment is to induct a system into this practice, not merely to push it closer to a label.
Human feedback must therefore exceed the binary. The signal should articulate correction: why a claim is ungrounded, which source repairs the gap, how structure clarifies, and where risk must be surfaced. The point is not to humiliate error but to situate it, revealing the grammar of repair. The model learns not simply that it is wrong, but how wrongness relates to the practice-why an omitted footnote matters, how an untested assumption fractures a valuation, why a silent risk violates fiduciary duty.
In this frame, safety is not an external guardrail bolted onto capability; it is judgment under uncertainty. A system that understands risk as regulatory, fiduciary, and reputational will treat uncertainty as an obligation-to disclose, to justify, to bound-rather than as a statistical curiosity. ZORA designs alignment signals that carry this grammar: domain‑specific rubrics, exemplar corrections, and structured rationales. The model sees not merely the desired output but the route by which the output earns acceptance.
To align is to induct. The decisive test is whether professionals adopt, amend, or reject the deliverable in their own workflow. When adoption becomes routine-when artifacts are used rather than admired- the signal has done its work. Intelligence then becomes responsible: answerable to a life with rules, capable of reasoning with evidence, speaking in structure, and acting with consequence. In short, alignment is not politeness; it is the education of a system into judgment.