Risk: self-modifying weights
Model weights are frozen
What changes is memory, notes, skills and the exam system. The model is a row in a catalog, swappable at any time. Growth assets stay.
Recursive Self-Improvement, with boundaries
It works for you by day, writes down what it learns on the spot, and consolidates memory at night. You set exams for its profession and score them. It polishes itself on a snapshot, and only a candidate that beats the baseline waits for your approval. One lap, one version, always reversible.
Access is by application. Leave your email and we send an activation email once approved.
Work Finishes the task on its own sandboxed computer and delivers files.
All present, none of them the point:
createrole is about what happens after the task ends. Does what it learned today still exist tomorrow? Will it do better next time?
Six stages. Each is a mechanism that actually runs in the product, with its own page.
Notes
Writes down what it learns while working
Three files on its drive, written by itself: who I am, my relationship with you, mistakes I have made. When it learns something about you, makes a promise or gets something wrong, it writes it down immediately. All three are in context every turn, capped at 1,500 characters each, enforced on write.
Consolidate
Distills the day into memory at night
One diary entry per day, transcribed per session in real time. In the early hours it distills the day into memory pages, rewrites its relationship notes and lessons, and updates what it knows about you. If nothing is worth keeping, nothing is written.
Night school
One click generates a syllabus and question bank by profession
It researches the profession in its soul profile online, and an examiner model writes a syllabus: skill dimensions by question type by difficulty. Calculation questions are verified by execution, knowledge questions must cite real sources, open questions follow a rubric. Practice and test pools are separate. Examiner and student use different models.
Eval
Scores the current version to set a baseline
A judge model scores each answer against the reference, with a one-line comment per question, giving the current version a number. The answering model never scores itself. Without a question set with reference answers, coaching does not start.
Coach
Polishes itself on a snapshot
Started by you. It scores the current soul first, then generates a candidate soul and scores it again. Only a candidate that strictly beats the baseline waits for your verdict. Anything lower is rejected automatically, and nothing changes before you decide. You see a score change, a rationale, and a line-by-line diff.
Version
You click apply, the version goes up
Every version pins the model, soul, skills and tools of that moment and says where it came from. Apply means v+1. Every soul revision is stored in full and can be restored. The next day it works with the new version.
Model weights stay frozen. What changes is memory, notes, skills and the exam system.
Session records answer what happened. The diary answers what the day looked like. Memory answers what to do next time. Recall has three layers: a few dated one-line hooks every turn, a full page from the terminal when needed, and only then the diary.
Lessons
The mistake, why, and what to do next time
Events
Episodes with lasting significance
People
Relationships, preferences, taboos, promises
Beliefs
Conclusions about the world and your business, with confidence
Procedures
How a given task gets done
Self
Who it thinks it is
Every one of these is a page on the trial platform.




A 1 minute 48 second recording: start coaching, wait for the candidate score, read the diff, click apply, and the version goes from v1 to v2. Waiting is sped up, nothing is cut.
Self-improving systems have two known risks: without an external anchor they rise and then collapse, and a judge that shares its source with the generator confirms itself. Each one is answered with structure, not a prompt.
Risk: self-modifying weights
What changes is memory, notes, skills and the exam system. The model is a row in a catalog, swappable at any time. Growth assets stay.
Risk: rewriting its own personality
Profession and style live in the database. There is no file for them on the drive, and the sandbox cannot see or write them. A coaching candidate goes live only when you click apply.
Risk: self-confirmation
The examiner runs on a separate, stronger model tier. The answering model never scores itself. The question bank never enters the employee's context.
Risk: getting worse
Every coaching run scores the baseline first. A candidate must score higher on the same set before you see it. Anything below is rejected automatically.
Risk: irreversibility
Every soul revision is stored as a full snapshot and can be restored. Nightly steps run once per day.
Risk: context bloat
The 1,500-character cap on notes is enforced on write. Deleting a conversation cleans up the memory and diary derived from it.
The loop, laid over one day.
By day it chats, plans, executes, delivers, and wakes up for due tasks. At night it consolidates memory, reviews relationships and lessons, and updates what it knows about you. Next morning, yesterday's lesson enters the first turn as a memory hook. Eval and coaching run when you start them.
One day of a digital employee
24hEvery step above is a mechanism that actually runs in the product.
Backend, frontend, sandbox service and memory service deploy to a customer's private network with one script. Unplug the network after acceptance and everything keeps working.

On two NVIDIA DGX Spark units, a department's entire AI headcount sits on a desk: one node runs local model inference, the other runs all application services, linked at 200GbE.
Research notes, method, releases.
Access is by application. Tell us what you want it to do, and we send an activation email once approved.