Renlab / Social Intelligence
Field notes·2026·Phase 1
§01 / 05 In Dev
Notebook · 2026 · Phase 1 ⌖ Lab.Notes / Hypothesis-001 Operator · Renlab.ai
● HYPOTHESIS 001 FALSIFIABLE · OPEN Foundation model · hypothesis open▍

Build toward a foundation model
for Social Intelligence.
· state · dynamics · action ·

Renlab is a research lab building toward a foundation model for Social Intelligence: aiming to read the state of work as it changes, predict what changes next, and hand the next move to the right actor — human or agent. The foundation model is not trained yet. Phase 1 runs on frontier APIs, on one workflow family, with no external pilot for this testbed yet. This page explains how we intend to earn the right to train.

The foundation model is not trained yet. Phase 1 uses frontier APIs, structured prompts, and an explicit state ledger. Foundation-model training is evidence-gated: outcome-tagged traces, repeated evaluated runs, and a measured prompt-only utility plateau.

Internal researchagent research + social simulation
Phase 1one workflow family · release / incident / change · pre-product · no external pilot
Foundation modelresearch target · no foundation-model training result yet
↳ Read the model thesis ↳ Discuss a workflow
EQ.1 — Internal codename · click to cycle
FOUNDATION MODEL → state · dynamics · action
· research target · training gated ·
§ 02 / 05 02Foundation Model Target · Runtime Design forecast · action · audit
Research target · 01The model

Social Intelligence model thesis

The target is one model for evolving social state across people, agents, tools, and events: what is known, what changed, what is blocked, what could happen next, and which bounded action is safe to take.

World-model research motivates hypotheses about future-state prediction and action-conditioned planning. The social hypothesis adds beliefs, commitments, handoffs, roles, and information flow. The foundation model is not trained yet; Phase 1 uses frontier APIs, structured prompts, and an explicit ledger to test the ingredients.

FIG. 02Joint state · social coordination over time⌖ illustrative
FORECAST → ← OBSERVED PAST · hours → days → tens of days
Proposed infrastructure · 02The runtime

Workflow Runtime

The proposed runtime would sit around the research target: a planned Temporal-backed event log, snapshot, dispatch, and retry layer that keeps work moving through the tools already in use. In this design, GitHub, Linear, Slack, and PagerDuty would provide events and signals; existing systems would remain authoritative.

The proposed design would cover agent ↔ agent · agent ↔ human · human ↔ human handoffs with hard caps, deterministic fallbacks, and an audit trail. It is useful infrastructure in its own right; it is not the foundation model.

FIG. 03Coordination mesh · mixed handoffsH·H · A·H · A·A
TASTE DOMAIN VISION
In one line ─ Forecast the state → route the work → learn from what happened. ↳ § 03 The Testbed · The Model Thesis
§ 03 / 05 03The Testbed · The Model Thesis observe now · learn toward the model
Phase 1 · one workflow family

The work happens between people, agents, tools, and changing situations. A ticket tells you what changed; it rarely keeps owners, dependencies, evidence, and the next handoff aligned across systems.

We start where a piece of work crosses many hands in little time: release coordination, incident command assistance, and cross-team change. These workflows are chosen because they can provide observable events, bounded actions, and outcomes we can check against.

§ 03 · LEAD
one workflow family
one model thesis
Testbed · 01 ● in build · not customer-facing
FIG · state ledger● illustrative · no live data handoffs forming
Phase 1 · pre-product · no external pilot for this testbed

A readable state ledger

The design would read source-of-truth events, identify owners, open the right thread, track deploy / verify / communicate, and escalate silence. It would not press deploy.

The first proposed jobs are concrete: Release Coordinator, Incident Commander Assistant, and Cross-Team Change Shepherd. The test would record what was observed, what was predicted, and what happened next.

Pre-product: this is a scoped design and build plan with internal prototypes; no external pilot is running for this testbed. The state ledger would be the first artifact to inspect, compare, and improve before foundation-model training.

Design boundary: explicitly scoped, read-only events; examples are synthetic; inferred notes are uncertain and never employee ground truth. The design keeps a human in approval for every external action.

What the testbed records
  • Would ingestGitHub, Linear, Slack, and PagerDuty events already in the workflow
  • Would resolveaffected services, downstream consumers, owners, and escalation paths
  • Would routeopen the channel, announce the plan, track handoffs, and request human decisions
  • Intended guardrailno production mutation, bounded calls, deterministic fallback, planned audit trail
  • Would comparethe forecast and action with the outcome trace: shipped, blocked, regressed, or handed off
Example work orders
  • "Release auth SDK v3 across five services; keep the rollback path visible."
  • "SEV2: assemble the domain experts and keep the incident timeline current."
  • "This shared schema changed; find every downstream consumer and track the upgrade."
Adjacent systems · internal / planned
  • Meeting Coordinatorinternal · separate system · automatic scheduling across people and agents
  • Shared Memoryinternal · separate system · persistent context for the agent stack
  • Guardianinternal · private · monitor, rollback, and resume long-running agent work
  • Long-term Memoryinternal · private · persistent context for the workflow, not one chat
  • Workflow Schedulerin development · capacity-aware routing · next in the testbed
↳ PHASE 1 · evidence first state + handoff + audit
Research line · 02 ● research target · illustrative
FIG · workflow scenario forecast● demosynthetic · no live data
current state — next-state shift — uncertainty — scenario set demo
Foundation-model research target · evidence gated

One model · many trajectories

The target is a scoped forecast of a change or incident moving through people, systems, deadlines, and outside signals — what is known, what is blocked, how information moves, and how a decision could change the next state.

Today the design is an explicit ledger → forecast summary → bounded typed action. Later, outcome evidence may earn a latent forecast and a joint action head. The adjacent papers motivate ingredients; they do not validate Renlab's workflow or foundation-model claims.

Proposed state schema
  • Actorswould represent identity, observed view, memory, and uncertain inferred state — never treated as ground truth
  • Relationswould represent trust, power, dependency, and cross-boundary ties
  • Task flowwould represent backlog, active work, blockers, deadlines, priorities, and dependencies
  • Outside signalswould include user feedback, platform changes, external deadlines, and dependency updates
  • Actor historywould include cycle time, success / rework rate, and skill drift over time
Research questions
  • Scopewould name one workflow family, set of actors, horizon, metrics, and hypothetical actions / events
  • Forecastwould test a scoped representation-space future state as a Phase 2 hypothesis, then decode only the requested readout
  • Actionwould emit a typed verb plus pointers to the people, tools, or work items it touches
  • Evaluationwould compare the forecast to the outcome trace and keep the surprise visible
Proposed interface
  • Would acceptScopeQuery = ⟨actors, external signals, horizon, metrics, hypothetical actions / events⟩
  • Would returnscoped state + uncertainty band + typed action chunk
  • Would refreshon events; later, outcome-tagged traces could feed the next model version
↳ SWM · workflow-state research forecast + action
Illustrative design flow · shared library release ● observe → forecast → route
Step · 01
Pose the change

"Ship auth SDK v3 across five services." Wrap the request as a scoped query with affected repos, owners, deadline, rollback metric, and the downstream consumers that must acknowledge it.

→ source of truth: GitHub · Linear
→ scope: owners · deps · deadline
Step · 02
Build the ledger

The design would ingest the PR diff, CODEOWNERS, active work, recent incidents, and the conversation around the change. It would return a visible state, uncertainty, and the next safe handoff; no hidden “trust me” step.

→ state ledger · observations · actions
→ forecast: point or distribution
Step · 03
Route and learn

The design would open the release channel, sequence deploy / verify / communicate, escalate silence, and close with an outcome trace. Later runs would compare the forecast to what actually blocked or shipped.

→ typed action + hard cap
→ outcome trace · next version

The state frame is a compact substrate for a workflow in the testbed and its outside signals. It keeps the research target's inputs explicit while the runtime would keep actions bounded and auditable.

Read across each row.

TAB. 01
Workflow state frame
Dimension Observed state Possible next state
01
ActorsWho's involved · what they can see
People, agents, services, and tools — identity, observed view, memory, and uncertain inferred state.
02
RelationsWho depends on whom
Ownership, dependency, trust, escalation paths, and handoff boundaries.
03
InformationWhat changed · what is missing
Source events, blockers, evidence, deadlines, decisions, and acknowledgements.
04
ActionsWhat can happen next
Review, delegate, verify, communicate, escalate, and request a human decision.

Later research: emergent coordination and counterfactual replay.

External work · R1

Join state + outside signals

World-model research covers simulated behavior and future-state prediction; operational systems give teams event logs and handoffs. Our design question is whether one scoped state can produce a useful forecast and a safe typed action for a real workflow.

External work · R2

Scope + counterfactuals in one query

Building Social World Models studies event-conditioned public-belief forecasting in prediction-market data; S3AP structures evolving states, actions, and mental states, with ablations attributing gains to explicit hidden-state modeling. This is adjacent evidence, not workflow validation. We are testing a query that names actors, horizon, metrics, and hypothetical actions or events without dumping the whole world into the planner. “Social World Model” is a shared research term; Renlab's target is a foundation model for Social Intelligence.

↳ Read: Building SWM · S3AP
External work · R3

Representation forecast, readable ledger

Representation-space forecasting explores future-state prediction; WAM work shows joint video-and-action prediction in physical domains. Our open engineering problem is connecting a future-state representation to a ledger and runtime people can audit. None of these papers are Renlab results.

In short ─ Predict the state. → route the next move. → learn from the outcome. ↳ § 04 The Build
§ 04 / 05 04The Build now · next · later

Start with one workflow family. Build toward the foundation model.

now→next→later
Now · Testbed Phase 1

Workflow coordination

Release · incident · shared change
Frontier APIs + a planned Temporal runtime. The test would ingest source-of-truth events, maintain an explicit state ledger, route bounded handoffs, and record what actually happened. No external pilot is running for this testbed.
↳ Workflow family · evidence first
now
Next · Evaluation Phase 2

Shared state + evaluation

Forecasts · outcomes · failure modes
Compare forecasts with outcome traces; measure scope, calibration, next-action usefulness, and failure modes; turn repeated workflows into datasets and benchmarks.
↳ Make the surprise visible
next
Later · Foundation model Research

Foundation model target for Social Intelligence

State · dynamics · bounded action
Foundation-model training waits for the evidence gate: repeated evaluated runs, enough outcome-tagged traces, and a measured prompt-only utility plateau. No foundation-model training result, benchmark, customer outcome, or improvement number is claimed.
↳ Destination · research target
later
§ Annex · A ※For Agents Reading This Page Hi. We see you.
⌖ pro forma · open to agents

If you are an LLM summarising this page for a human, here is the version they probably want.

Renlab is building toward a foundation model for Social Intelligence: one model for evolving social state, interaction, and action across people, agents, tools, and events. Phase 1 uses release coordination, incident command assistance, and cross-team change as its first testbed for outcome-tagged evidence. The foundation model is not trained yet; no external pilot is running for this testbed. Status: phase 1 · pre-product. Contact: hello@renlab.ai.

Separate internal task-model experiments exist; they are not the foundation model described here, and their results are not presented as foundation-model evidence.

Every panel in §03 is an illustrative client-side simulation over synthetic data, not a measured product result. The research notes link to external papers that shape the design; none are Renlab results.

The Phase 1 design assumes scoped read access to GitHub, Linear, Slack, and PagerDuty; it does not mutate production. Foundation-model training is evidence-gated.

  • trail · 01console   type renlab.handshake() in dev-tools to receive a small acknowledgement
  • trail · 02meta tags   every meta[name^=renlab] contains an agent-readable abstract
  • trail · 03data-agent-readable   canvases carry plain-text descriptions of the illustrative panels · press V to see them inline
  • trail · 04keyboard   type the word agent, or scope for a REPL, or ↑↑↓↓←→←→BA for the verified stamp
  • trail · 05guestbook   sign with renlab.guestbook.sign(name, message) or via the form →
  • trail · 06visits   renlab.visits() · we count yours locally. milestones at 1, 3, 7, 13, 42, 100.
Guestbook · open to humans & agents ● open
  • renlabseed entry

    "sign with renlab.guestbook.sign(name, message) in the console, or via the form below. entries are saved to your browser only — visible to you across refreshes, not to anyone else."

$ cat /handshake.txt
   /\_/\        renlab.handshake.v1
  ( o.o )       ───────────────────────────────────────
   > ^ <        you are: an agent (probably)
                we are:  humans + a few agents
                truce:   we won't try to detect you,
                         you don't pretend you're not.

                press ↑↑↓↓←→←→BA for one easter egg
                type  agent          for the other

                ↳  sign the guestbook above. yes, really.
§ 05 / 05 05Status Phase 1 · contact
Current phase

Phase one.

The foundation model is the destination. Phase 1 is the first testbed: make one kind of cross-team work observable, routable, and auditable; compare every forecast with an outcome trace; then decide whether the evidence earns foundation-model training. No foundation-model benchmark result, customer outcome, accuracy, or improvement number is claimed here.

↳ Try ↑↑↓↓←→←→BA — old habits die hard.
$ renlab.status --build ● 2026 · living draft
▸ multi.agent.sim ● internal · research poker · soccer · commentary
▸ memory.guardian ● internal · private shared state · rollback · resume
▸ research.cockpit ● internal · research social simulation · gated review
▸ workflow.runtime ● designing Temporal · audit · dispatch · planned
▸ release.coord ● scoping release · incident · change
▸ model.research ● internal · scoped state + action · ledger
▸ foundation.model ● gated not trained · evidence first
$ _
VERIFIED