AI4C2 · research foundations

Beyond the Model

Harnesses and Architectures for Language Agents

Charles F. Vardeman II
Center for Research Computing, University of Notre Dame
August 18, 2026
Start here · agent-assisted walkthrough

Explore this deck with ChatGPT or Codex

Copy and paste into ChatGPT or Codex
Open in your built-in browser if available, or use the HTML I attach:
https://la3d.github.io/nuggets/projects/ai4c2/updates/2026-08-18/
Load per-slide Web Annotation JSON-LD from raw HTML. Confirm success or ask for the HTML.
Use visible context or my slide number/title; ask only if ambiguous. Ground answers
in the actual slide, diagrams, and annotation. Be brief, conversational, and plain;
expand when I ask. I control navigation: no unsolicited lectures, advancement, or quizzes.
For deeper questions, follow relevant linked original papers or YouTube references.
Distinguish slide claims, source-supported findings, added explanation, and uncertainty.
For videos, disclose access to transcript, video content, or metadata only.
Never imply you watched content you could not access.
Research foundations

From next-token prediction to scientific investigation

What changes when an agent can operate on knowledge and methods outside its immediate context?

1 · Model foundationsPrediction, training, and learning from context
2 · Agents and harnessesExecution, memory, skills, and feedback
3 · External workspacesRLMs, domain objects, and explicit operations

Research destination: test whether this improves correctness, reuse, and transfer.

Language models · one token at a time

Each next token depends on the context so far.

The context changes with every generated token—so the next prediction changes, too.

Context tokens flow through attention and model scoring to next-token probabilities; the selected token is appended and the extended context returns for another prediction. All example token splits and probabilities are illustrative, not measured.

Attention weights mix information within the model; they are not the final next-token probabilities.

Explore: 3Blue1Brown · Attention, step by step · Transformer Explainer · interactive model

Research foundations

Training changes weights; inference uses them

Training

Adjust learned numerical parameters—weights—using examples and a training objective.

Inference

Use those weights and the current context to generate an output.

Training accuracy: performance on examples used for learning.

Held-out accuracy: performance on examples excluded from training.

Doing well on familiar examples is different from generalizing to new ones.

Training dynamics

Memorization and generalization can arrive at different times

Animated grokking example: training accuracy rises quickly while test accuracy remains near chance, then test accuracy rises sharply after prolonged training

Animation: Pearce et al., “Do Machine Learning Models Memorize or Generalize?” · phenomenon: Power et al., “Grokking.”

Inference-time adaptation

Examples in the prompt teach a temporary task

Task · sentiment classification

Decide whether a review expresses a positive or negative opinion.

PROMPT

Classify each review as POSITIVE or NEGATIVE.

I loved this movie.POSITIVE
This was terrible.NEGATIVE
Surprisingly enjoyable.?
→
MODEL COMPLETION POSITIVE

No retraining: the examples shape behavior only while they remain in the prompt.

In-context learning: Brown et al., “Language Models are Few-Shot Learners” · Garg et al., “What Can Transformers Learn In-Context?”

Inside the transformer

An induction head finds a pattern and continues it

An attention head is a small part of a transformer that decides which earlier words are useful for predicting the next one.

EARLIER IN THE PROMPT The code word is BLUEBIRD.
look back at what followed the same words
LATER IN THE PROMPT The code word is → BLUEBIRD
1Find a familiar sequence
2See what followed it earlier
3Use that continuation again

Mechanistic evidence: Olsson et al., “In-context Learning and Induction Heads.”

Continue learning

Three visual explanations worth watching

All three videos: Welch Labs · companion grokking animation

Further viewing

Build the full mental model—from tokens to ChatGPT

Video thumbnail showing Andrej Karpathy beside a visualization of the computations inside ChatGPT
Andrej Karpathy · recommended introduction

Deep Dive into LLMs like ChatGPT

  • How text becomes tokens and next-token predictions
  • How pretraining and post-training shape behavior
  • Why tools and agent loops extend the base model

Watch: “Deep Dive into LLMs like ChatGPT” · long-form visual introduction.

From prediction to tool use

Tool calls request execution outside the model.

A calculator tool can be an ordinary local Python function—no HTTP request required.

The model sees the add tool contract and generates add with arguments 2 and 3. The harness validates and dispatches to a local Python interpreter; the function returns 5, which joins context so the model can continue. Source code stays in the execution environment.

Weather API: the same request–execution–result loop, with a remote service instead of local arithmetic.

Tool semantics: Anthropic · Tool use overview · Handling calls and results · calculator example is schematic.

From generation to action

An agent puts tools in a loop

“An LLM agent runs tools in a loop to achieve a goal.”

1 · Goal

A bounded objective and a condition for stopping.

2 · Tools

Actions that retrieve information or change an external environment.

3 · Feedback

Tool results return as observations that shape the next model call.

Definition: Simon Willison, “I think ‘agent’ may finally have a widely enough agreed upon definition.”

Anthropic diagram comparing single-turn prompt engineering with iterative context engineering: candidate documents, tools, memory, instructions, and history are curated into selected context; the model generates an assistant message or tool call, and tool results return to the candidate information pool.

Source: Anthropic, Effective context engineering for AI agents (2025) · original diagram.

Harness responsibilities

The harness manages what reaches the next model call.

Execute tools and capture results; maintain files, memory, and intermediate work; assemble selected context. A larger working environment feeds a smaller active context window through harness curation, then the model. Stored information is not automatically loaded.

The model can choose retrieval and tool calls; the harness implements the mechanisms and policies.

Design synthesis: Anthropic · Context engineering · Zhou et al. · Externalization, §6.

Multi-step reinforcement learning

Agentic RL trains over trajectories—not single responses

Policy actions pass through the harness to the environment; observations return and completed trajectories train the policy. Task actions calculate/query/inspect; continuity actions save findings/retrieve evidence/update plans. These purpose-based action examples are the deck’s teaching synthesis.

Having tools differs from learning to use them well. Evaluate whether goals and evidence survive across steps.

Loop foundation: Microsoft, MAI-Thinking-1, Fig. 18 / §3.3, p. 40 · continuity examples added here, not Microsoft implementation claims.

What is being learned?

Distinguish three RL learning targets

Model policy

Train through the deployment harness.

The harness constructs context; the model learns its responses and tool requests.

Agent Lightning v1.0: summarization and subagents can break token-prefix continuity.

Memory decisions

Learn what to archive and revisit.

Write/read, compress/retrieve: train behavior with task reward under a context budget.

MemexRL: indexed summaries plus recoverable external experience.

Harness controller

Choose operations around a frozen executor.

Observe · retrieve · call-tool · draft · check · revise · submit.

Offline RL: select the next operation; no executor weight or prompt update.

A capability is available machinery; a learned policy decides how to use it. Memory behavior can be part of the model policy.

Agent Lightning v1.0 · MemexRL · Offline harness-control RL · separate research settings, not a single combined system.

Evaluate long trajectories

Evaluate outcomes and continuity

METR · task-completion horizon

Human-expert task duration
at a stated success probability
50% or 80%

Measures the evaluated agent system on a task suite.

Not the model’s wall-clock runtime, a universal coherence score, or an isolation of RL’s effect.

LOOP / AppWorld · training study

Train in the environment
then test on held-out tasks.

Reports stronger task performance, documentation consultation, and recovery behavior.

Inspect matched-base and training-method comparisons; a cross-model ranking is not a causal ablation.

Our evaluation question: did goals, evidence, decisions, and progress survive the trajectory?
Pair task outcomes with trace checks, recovery tests, and context/cost budgets.

METR · definition and methodology · Apple · LOOP · Chen et al., §5 / App. A, E.

Define cognitive architecture · CoALA, 2023/2024

CoALA: organizing a language agent as cognitive architecture

A cognitive architecture specifies how memory, reasoning, action, and decision-making work together within an agent.

CoALA organizes memory storage, internal and external actions, and a repeating decision procedure.

Definition: our synthesis. Diagram adapted from Sumers et al., CoALA, Figs. 4–5 and §4. A conceptual framework for language agents.

From functions to implementation choices

From cognitive functions to externalized machinery

CoALA · functional organization

What works together?

Memory and internal actions
Grounding in an environment
A coordinating decision cycle

Externalization · systems emphasis

How is it realized?

Persistent state and skill artifacts
Explicit interaction interfaces
Harness control and feedback

Trace a function into its artifacts, interfaces, and runtime mechanisms.
A change of emphasis; the frameworks do not map module for module.

Zhou et al., §1–2, 6–7 names CoALA its closest conceptual bridge. Comparison is our synthesis.

Externalization · placement

Capability spans weights, context, and runtime

Weights

Learned language and reasoning capabilities

Context

Selected evidence and guidance for this call

Runtime

Persistent artifacts, tools, and control

Externalization changes the work presented to the model.
Retrieve relevant state · reuse procedures · invoke structured operations

Zhou et al., §2 · Teaching summary.

Memory types · different jobs

Memory types preserve different kinds of value

Working · stay on task

Active goals, plans, hypotheses, and intermediate results.

Episodic · learn from events

What happened, under which conditions, and with what outcome.

Semantic · reuse knowledge

Facts, relationships, and consolidated understanding.

Procedural · reuse methods

How to act: routines, strategies, and skills.

Type guides retention, retrieval, and revision.
Personalization can apply across types; these are not four isolated databases.

CoALA, §4 · Zhou et al., §3 · Taxonomy synthesis.

Memory lifecycle · WikiSkill example

Preserve experience; consolidate what it teaches

1 · Retain traces

Keep the original execution evidence.

2 · Maintain a wiki

Organize patterns, workarounds, and evolution history.

3 · Inform a proposal

Use accumulated knowledge to guide a skill change.

Our memory lens: episodic evidence → consolidated knowledge
The wiki mixes abstractions and history; it is not purely semantic memory.

Tang et al., WikiSkill, §3 / Fig. 2 · Memory interpretation is ours.

Procedural expertise · WikiSkill continued

Reusable skills need an acceptance test

Package the method

Steps, decision rules, constraints, and checks.

Test the change

Evaluate a proposed skill on validation tasks.

Keep or roll back

Retain improvements; reject unsuccessful edits.

WikiSkill retains the wiki when a skill edit is rejected.
Task runs use active skills → new traces feed the next evolution cycle.

Zhou et al., §4 · WikiSkill, §3.2 / Algorithm 1.

Interaction structure · teaching example

A tool call needs more than argument types

Request

search_reports(query, filters)
Discover the operation and validate inputs.

Execution

Check permission; track completion, timeout, or failure.

Result

Return evidence references or a structured error.

Protocols specify the exchange; runtime mechanisms enforce it.
Tool = operation · skill = guidance for use · protocol = interaction contract

Zhou et al., §5 · Report-search example is ours.

Harness design · coordinated execution

The harness makes the pieces work together

Select and budget

Which memory and skill content enters this call?

Permit and execute

Which actions may run, and when must a person review?

Observe and recover

What happened, what failed, and what runs next?

Memory + skills + protocols operate inside these controls.
Same model, different runtime choices, different operating conditions.

Zhou et al., §6.1–6.4 · Three-question grouping is ours.

System synthesis · report-analysis example

One task connects memory, skills, and feedback

Before the call

Retrieve prior findings.
Load the analysis procedure.

During execution

Issue a structured search.
Check and interpret its result.

After the result

Update the active plan.
Save evidence and the outcome.

Across runs: examine failures → propose a better procedure → test it
WikiSkill makes this slower evolution loop concrete; the report example is our synthesis.

Design question: what should persist externally, and what will selection, loading, and maintenance cost?

Zhou et al., §7.1–7.3 · WikiSkill · Example is ours.

Recursive Language Models

An investigation can exceed the context window

A report archive, ontology, or inferred graph can exceed the controller’s context window. Keep the full objects external; inspect them through code and bounded subagent questions.

Store it as data

Keep the full archive outside the model’s immediate conversation.

Search before reading

Inspect and select the portions relevant to the present question.

Ask smaller questions

Solve bounded subproblems and combine their results in code.

Source: Zhang, Kraska, and Khattab, “Recursive Language Models,” v3.

RLM mechanism · following Madura’s introduction

An RLM operates on context as data

Symbols refer to objects
Reference large objects without copying them into context.

Subcalls are functions
Delegate bounded inspection; return selected findings.

The model chooses the path
Code filters and computes; the controller sees selected evidence.

Original RLM diagram: the root model accesses context through an environment, calls submodels or nested RLMs, and receives their results back into the environment.

Original diagram: Zhang and Khattab, RLM introduction, Fig. 1 · Madura, “It’s Tokens All The Way Down”.

RLM evidence · scope matters

A small experiment tests document scaling

Original two-panel chart comparing answer rate and API cost as document count increases.

20 sampled questions · 10–1,000 documents · supporting evidence included.
Perfect performance at 1,000 documents applies to this subset, not the full benchmark.

Original chart: Zhang, Fig. 5 · Benchmark: Chen et al., BrowseComp-Plus. Historical API costs.

Continual adaptation · external state

Prime Agent learns through harness refinement

Observe experience

Retain trajectories and feedback.

Refine artifacts

Revise memories, skills, prompt notes, and agent roles.

Change later behavior

Load selected updates into subsequent calls.

Persistent RLM workspace + versioned, revisable harness state
Model weights stay fixed during this refinement loop.

Karten et al., Prime Agent, §2.2 / §2.5 · Connection to WikiSkill is our synthesis.

Harvey · harness comparison

Review a data room; produce a diligence memo

Data room: a company’s document collection for an acquisition review.
Hard part: connect scattered evidence, notice missing support, and produce a cited assessment.

Harvey RLM architecture: a data room is loaded into a root agent’s Python REPL as queryable variables; the root delegates batches to subagents, receives findings, and synthesizes a memo. Only printed outputs enter the root context.

Tool-loop → RLM harness: 23.3% → 62.4%
Mean rubric criteria pass rate across seven models · 50 held-out synthetic data rooms.

Original diagram, Fig. 5, and reported results: Harvey, “Post-Training RLM Agents for End-to-End M&A Diligence.”

Harvey · matched root-model RL comparison

Training improves root coordination

Task quality

29.9% → 63.0%

Rubric criteria pass rate
on held-out data rooms

Evidence coverage

62% → 96%

Measured content coverage
without a coverage reward

Work organization

Delegate and synthesize

More subagent calls;
memo writing interleaved
with review of findings

Train the root with RL; hold subagent models fixed.
Matched base and trained checkpoint · 50 held-out synthetic data rooms · reported by Harvey.

Our next question: better coordination assembles evidence—
how do we check the claims derived from it?

Harvey, Reinforcement Learning · Fig. 12–13 / Table 2. Behavioral changes accompany the gain; they do not isolate its cause.

Neurosymbolic motivation

Why agentic systems need ontologies

Video thumbnail featuring Frank Coyle with the words Ontologies Keep Agents Honest
Frank Coyle · UC Berkeley · AI Engineer

Probabilistic reasoning inside.
Logical guardrails outside.

  • Combine learned interpretation with explicit knowledge.
  • Represent domain meaning outside the model.
  • Check intermediate results within the agent loop.
  • Our extension: derive consequences as well as check proposals.

Watch: “Why Agentic Systems Need Ontologies” · 21 minutes.

Reasoning vocabulary · complementary roles

Induce patterns, propose explanations, derive predictions

Induction · learn a pattern
Across repeated trials, heating this rod accompanies expansion.
↓
Heating tends to expand this rod.
Abduction · propose an explanation
The rod has expanded. What could explain it?
↓
Perhaps it was heated.
Deduction · derive a consequence
Given: heating expands this rod under specified conditions.Suppose it is heated under those conditions.
↓
It will expand.

Investigation loop: observe → propose explanations → derive predictions → check evidence → revise.

General architecture · proposed skill and memory integration

Retain knowledge—and methods for using it

Domain knowledge

Data · ontologies · rules
What concepts mean and relate to

Procedural knowledge

Skills · query templates · checks
How to investigate

Experience

Findings · attempts · failures
What previous work taught us

Named, versioned objects · discover and inspect selectively
↓

Agent proposes hypotheses, selects inputs, executes operations
Retain results + traces → inspect evidence → choose the next step

Across runs: experience → proposed skill revision → evaluate → accept / roll back

WikiSkill + Prime Agent motivate this synthesis. Full integration is proposed; the next slide shows our implemented substrate.
Our workspace prototype · research implementation

Investigate through retained objects

Large ontology and data objects feed a query and retained result. Selected evidence returns to the controller; a dashed broader-RLM branch delegates bounded inspection to a subagent.

Implementation: Linked Science Cloud · Loading an ontology does not itself perform inference.

Our workspace prototype · linked data across sources

Shared identifiers connect separate sources

CONFIGURED FEDERATION · UNIPROT–RHEA–WIKIDATA
UniProt

Proteins
Catalytic activities
Protein accessions

Rhea

Reactions
Reaction participants
Reaction identifiers

Wikidata

Linked entities
UniProt accession mapping
Additional relations

Join points: UniProt ↔ Rhea reaction references
UniProt ↔ Wikidata accession mapping

SEPARATE READ PATH

WikiPathways

Pathway records
Source identifiers

Further joins need explicit, checked mappings.
Loaded knowledge

Ontologies + mappings
PEEK locates objects

→
SPARQL through the REPL

Query selected sources
Join on explicit identifiers

→
Retained result

Rows or graph + provenance
Inspect · reuse · analyze

↓ Source responses feed SPARQL. Full results stay external; selected views return to the model. Research prototype.

Worked example · proposed reasoning integration

Separate supplied facts from derived facts

REPL object Small, inspectable example
Base graph A: 5 miles, 8 kilometers · B: 10 miles
Rule set Compute kilometers only when the base graph lacks that value
Inference graph B: 16 kilometers, using the example’s integer conversion

Why separate them? Preserve A’s supplied value; identify B’s value as computed. Later operations can choose original data or derived data.

Illustration from SPARQL 1.2 RL §3.9 · W3C Working Draft · object integration is proposed.

Ontology and instance · two graph objects

A reusable pattern gives observations structure

Paired graphs: the reusable SOSA observation pattern links an Observation to Sensor, Observable property, Feature of interest, and Result. A matching instance links Observation 42 to Sensor 7, Water level, Lock Gate 3, and 314 centimeters. Separate ontologyGraph and observationGraph handles identify the two retained objects.

Left: conceptual ontology view, not instance triples. Right: illustrative data using those relations. Both graphs can be inspected and passed to operations.

Teaching adaptation: W3C SOSA/SSN · Colpaert’s water-level example.

Knowledge engineering · sensor example

One observation, two representations

SOURCE · LOCAL VOCABULARY

314 cm

Water level at lock gate 3
Sensor 7 · same observation time

→
APPLICATION · SOSA PROFILE

3.14 m

Water level at lock gate 3
Sensor 7 · same observation time

Source provides

Fields and units
SHACL input shape

Knowledge connects

Meaning and conversion
Vocabulary alignments + unit rules

Application requires

Accepted representation
SHACL output shape

SOSA provides a reusable observation pattern; alignments connect local terms to it. Colpaert · water-level example.

Executable knowledge · prepare once, reuse

Compile the plan; reuse it across observations

PREPARE · WHEN THE CONTRACTS CHANGE

Source shape + accepted alignments + target shape
↓
Compact processing plan

Incoming messages

314 cm
298 cm
402 cm

→
Reuse the plan

Align terms
Convert units
Select output

→
Application views

3.14 m
2.98 m
4.02 m

RLM connection: find, propose, inspect, and reuse these artifacts.
Accepted rules perform the repeated transformation.

Planning assumes the source contract holds; validate separately when uncertain. Colpaert · shape planning.

Beyond the exemplar

Different procedures, explicit result types

Operation Inputs → retained output
Abductive reasoning Observations + rules → candidate explanations
Rule inference Facts + rules → derived facts
Constraint checking Data + constraints → validation report
Theorem proving Premises + goal → proof or solver status
Numerical computation Values + algorithm → computed result
Simulation Model + assumptions → simulated outcomes

Shared architecture: select objects → execute → retain → inspect. Guarantees depend on the procedure, its inputs, and its assumptions.

Research foundations

AI can help build the knowledge it investigates

DisMech: an ontology-grounded collection of disease mechanisms, assembled with AI-assisted curation.

Represent

Structured disease records
Ontology terms
Mechanistic relationships

Support

Evidence quotations
Source references
Hypothesis status

Check

Schema and term checks
Reference checks
Model-assisted review

Our proposed extension: investigate these artifacts as retained workspace objects; query, compare, and reuse intermediate findings.

DisMech methodology · Repository · Its checks do not establish scientific truth.
Why this might generalize · Zhang and Khattab

Different tasks can reuse the same strategy

Task A · classify questions

Group question types and count them.

Task B · classify spam

Group spam / non-spam and count them.

↓   Shared strategy   ↓

Partition → classify bounded pieces → aggregate

Keep task-specific data and intermediate results in REPL objects.
The root model coordinates a familiar sequence.

Reported: transfer across longer tasks and domains sharing a strategy.
Our hypothesis: explicit domain knowledge makes more operations reusable.

Zhang & Khattab, Figs. 6–8 · schematic task pair adapted from their OOLONG transfer experiment.

Research hypothesis

An agent can work with explicit knowledge

Investigate evidence, apply domain knowledge and reusable skills, and retain results and experience for checking and reuse.

Harvey

Investigate an external collection

Our workspace prototype

Data and ontologies as objects
Research prototype

Colpaert

Execute reusable domain knowledge

DisMech

Build structured knowledge with AI

Hypothesis: this combination improves correctness, traceability, reuse, and transfer to new tasks.

Test 1: add explicit knowledge and reasoning to a matched RLM.
Test 2: test evaluated skill refinement across runs against fixed skills.
Compare correctness, traceability, reuse and transfer at comparable budgets; report failures and cost.

Synthesis and proposed evaluation · DisMech · preceding sources provide components, not validation of the combined system.

Continue the conversation

Questions & discussion

What should an agent retain, reuse, and learn outside the model?

Data · ontologies · skills · experience

Explore the deck   ·   Prototype repository

AI4C2 · Agentic AI for C2 Platforms