Proven applications, maintained adaptations
Recommendation: Trellis should make good existing software your own—and keep it working after you change it. For the first proof, keep Memos and Plane, one operating experience, one replaceable agent, and then one useful adaptation that survives an upgrade. Ambition belongs in the continuing promise, not the number of components.
Origin: the founder’s latest direction (repository: docs/source/2026-09-08-founder-malleable-direction.md), following the linked research below. Decision: assistant recommendation, not an accepted expansion of release scope. Evidence: primary-source reading on 2026-09-08 plus read-only product critique; no new runtime capability was implemented or tested for this paper.
Subsequent analysis, September 10: The stewardship-contract paper (repository: docs/research/2026-09-10-stewardship-contract-paper.md) adds historical provenance, a motivated contract, alternatives and a worked Plane application. The August contracts predate this reading ledger; these sources support subsequent refinement, not a reconstructed origin story for the original schema.
What was actually reviewed
This revisits the thirteen linked essays, paper, architecture/specification pages and book chapter behind the preceding product discussion, plus Entire’s checkpoint documentation and installed behavior. It is not a claim to have read every cited book, linked bibliography, source repository or research implementation. Each entry states its reading scope. Technical specifications and research prototypes provide ideas and constraints, not automatic proof that their guarantees transfer to Trellis.
R01 — Malleable software: start with something worth keeping
Malleable software, Geoffrey Litt, Josh Horowitz, Peter van Hardenberg and Todd Matthews, 2025. Read the essay’s user-agency argument, design principles and prototype discussion.
Reading: customization should grow from using existing tools, not require every user to become the author of a replacement application. Its communal-creation argument is as important as AI-assisted programming: useful modifications can travel between people. Shared data and adaptable interfaces are ambitions, while cohesive existing applications still serve real purposes.
Our inference: Trellis should maintain a small, inspectable difference from an upstream application. Configuration and supported extensions come before a fork. A modification should explain its purpose and remain removable and shareable.
Tension/test: generating a patch is cheap compared with maintaining compatibility. The demonstration is not “AI changed the UI”; it is “the change still works after an upstream upgrade, or Trellis identifies why it cannot safely continue.” Research prototypes are not a ready-made maintenance platform.
R02 — Local-first: ownership is more than the server address
Local-first software, Martin Kleppmann, Adam Wiggins, Peter van Hardenberg and Mark McGranaghan, 2019. Read the seven ideals, technology comparison, prototype findings and future-work discussion; the product comparisons are historical.
Reading: user-device data, usable offline work, collaboration, longevity and control are distinct goals. Servers can support synchronization, storage and computation without owning the only useful copy. The paper also surfaces unresolved user-interface, permissions and history problems.
Our inference: Trellis needs exportable definitions, tested recovery and an exit that survives Trellis disappearing. Placement profiles should separately describe application storage, model processing, backups and dependencies.
Tension/test: hosting Plane yourself does not turn its web client into a local-first application. Local, hybrid and hosted Trellis must be honest deployment profiles, not promises of identical offline behavior. Test that applications and native recovery remain usable without Trellis or its model; do not make CRDT retrofits a release dependency.
R03 — Mixed initiative: uncertainty should change the interaction
Principles of Mixed-Initiative User Interfaces, Eric Horvitz, CHI 1999. Followed the publication page to the paper; read principles, LookOut, uncertainty, timing and learning sections.
Reading: useful automation combines user control with assistance sensitive to uncertain intent and interruption cost. LookOut works inside an existing calendar/mail interaction; it does not require replacing all direct manipulation with an agent. Asking is itself an action with benefits and costs.
Our inference: routine authorized checks should run quietly; uncertain, consequential choices deserve a concrete proposal. Keep direct controls for repeatable actions alongside conversation.
Tension/test: confidence about what a person wants is not permission to act. The approval boundary remains explicit. Test a mistaken request, cancellation, deferred choice and agent failure—not just a successful natural-language command. Do not introduce an attention-surveillance model merely because the research explored one.
R04 — Calm technology: quiet is not invisible
The Coming Age of Calm Technology, Mark Weiser and John Seely Brown, 1996. Read the essay, including center/periphery and situated-awareness examples.
Reading: information can remain available without constantly occupying focal attention; users should be able to bring peripheral information forward when needed. This is not an argument for withholding state or forcing awareness on everyone nearby.
Our inference: the Trellis home screen should show running apps and a small decision queue. A backup problem gets an actionable exception; successful routine checks remain inspectable history. No compulsory daily briefing.
Tension/test: fewer notifications can conceal missed incidents. Evaluate useful interruptions, missed important events and recovery time together. Make stale observations visible. A quiet system with unknown recovery is not calm—it is opaque.
R05 — Juju: package operational knowledge, not an imaginary employee
Juju and Charms architecture, Canonical. Read the workload, charm, client, controller, agent and Pebble architecture sections.
Reading: a charm encodes application-aware operational behavior in an established orchestration model. Juju’s agents and hooks are operational machinery, not necessarily language models. Its integration model is a useful precedent for relationships between applications.
Our inference: keep the existing App Definition/Steward Pack as a portable bundle of tested knowledge and referenced deployment assets. Reuse upstream operational assets where they fit. Agent instructions help discover and explain; reviewed operations do the work.
Tension/test: a Juju charm is not a standalone script you can transplant without its runtime assumptions. Do not adopt Juju, Kubernetes and a custom operator framework simultaneously. The test is adding a second application without branching the core around its name, while still encoding its actual backup and upgrade semantics.
R06 — Cambria: connections preserve meaning, not merely field names
Project Cambria: Translate your data with lenses, Geoffrey Litt, Peter van Hardenberg and Orion Henry, 2020. Followed the short project page to the essay; read schema evolution, lens workflow, findings and data-augmentation limitations. This is not an implementation audit of its appendices.
Reading: translations can help versions and applications cooperate, but some desired properties conflict. A single-assignee interface cannot faithfully expose every operation on a multi-assignee record. Missing related data is not magically reconstructed by changing a schema.
Our inference: start with one explicit, versioned connection whose direction, identity mapping and failure policy are known. Keep source applications authoritative.
Tension/test: a universal context graph must not hide semantic loss or expand permissions. Test deletion, duplicate delivery, unavailable lookup and incompatible schemas for the first connection. Sometimes read-only linking or a one-way operation is the correct product. Do not select Cambria as a production dependency on the strength of its conceptual fit.
R07 — UCAN: transferable authority still needs an enforcing endpoint
UCAN Delegation Specification. Read payload, subject/resource semantics, commands/policies and token-validation requirements. Not a library security audit or a review of every related specification.
Reading: delegated capabilities describe constrained authority across principals; execution-time validity and chain alignment matter. Signed metadata is explicitly distinct from delegated authority. An executor must understand what a resource and command mean.
Our inference: importing a customization imports requested capabilities, never credentials or permission. A new owner must grant authority for their own instance. Instructions and provenance remain data.
Tension/test: signatures cannot turn a broad Docker socket into a narrow API. Resource enforcement still needs engineering. Test that a shared modification cannot target another instance, widen an operation or reuse expired authorization. Learn from capability security without introducing a new token language into the first release.
R08 — Temporal: resume work, not a conversation
Workflow Execution, Temporal. Read the execution model, history, commands, states, execution chains and references to retries; not the linked retry documentation, full SDK or failure-semantics documentation.
Reading: durable execution records progress so work can continue across worker failures; deterministic workflow behavior and external activities play different roles. Conversation memory is not that execution history.
Our inference: an operation must survive the proposing agent disappearing. Persist enough state to inspect, resume or refuse safely, and re-observe after uncertain external effects.
Tension/test: durable orchestration does not grant arbitrary external effects exactly-once semantics. A process that dies after issuing a restart needs observation before retry. Preserve the current durable operation work; add Temporal only if demonstrated long-running coordination makes it simpler than the existing design.
R09 — Scaling agents: coordination is a cost, not a feature
Towards a science of scaling agent systems, Yubin Kim and Xin Liu, Google Research, 2026. Read the authors’ research article, methods summary, task comparisons and limitations; not a replication or full review of the linked paper.
Reading: the evaluated tasks respond differently to multi-agent architectures. Parallelizable work can benefit while sequential work can lose performance to coordination and interference. The results are benchmark-specific, not a universal agent-count rule.
Our inference: application expertise belongs in reusable definitions, not mandatory always-running agents. Use one replaceable agent for normal interaction; delegate independent research or review when it has a clear outcome.
Tension/test: measure completed work, user corrections, latency and model cost. More agents and more tokens are not evidence of progress. A second agent must earn its place through a concrete quality or speed improvement.
R10 — Beyond chat: use language to change the tool
Is Chat a Good UI for AI?, Geoffrey Litt, 2025. Read the article’s examples and interaction argument. Bibliographic title corrected on September 10; the earlier link label was descriptive, not the article’s title.
Reading: natural language and conventional interfaces complement each other. Repeated work can become a persistent interface instead of a repeated conversation. Existing applications can be valuable starting points rather than obstacles to replace.
Our inference: let someone request a change conversationally, then retain it as a usable setting, workflow, button or template in the application. Keep the Trellis operations interface small and explicit.
Tension/test: an impressive chat demo can leave the user with more work tomorrow. After the adaptation, ask the user to do the task again without prompting an agent. If it still needs a long conversation, we may have automated a demonstration rather than improved the tool.
R11 — Kubernetes operators: encode the application’s hard parts
Operator pattern, Kubernetes documentation. Read the operator concept and operational examples, not the navigation tree or a particular operator’s code.
Reading: a controller combines desired state with domain knowledge about lifecycle behavior. The examples extend beyond starting a container to backup, upgrades and application-specific recovery.
Our inference: the value of a Trellis definition lies in what it knows about this application’s persistence, versions and failure modes. A generic health endpoint and a restart command are insufficient stewardship.
Tension/test: naming something an operator adds no reliability. Restore real application content into a disposable environment and verify it. Borrow the pattern without forcing a small company to run Kubernetes.
R12 — in-toto: provenance has a declared coverage boundary
in-toto getting started, in-toto project. Read the example workflow, materials/products, layout and verification rules; not a complete supply-chain threat-model review.
Reading: provenance relates declared inputs, outputs, steps and authorized actors. The example makes coverage important: undeclared or weakly constrained artifacts do not become trustworthy because a signed record exists.
Our inference: keep the chain from request and upstream source to modification, review, test and released revision. Record source reuse and licensing separately from conceptual influence.
Tension/test: Entire captures conversation history; it does not prove the code is correct, every influence was captured, or the founder approved deployment. Start with linked Markdown, Git and actual checkpoints. Require a reader to trace one changed behavior back to its source and evidence before adding an attestation platform.
R13 — SRE toil: ownership must not become a second job
Eliminating Toil, Vivek Rau, edited by Betsy Beyer, in Google’s SRE book. Read the chapter, not the whole book.
Reading: repetitive manual service work that creates no lasting improvement is different from engineering. Even invoking an automation manually can remain toil. Some operational work remains necessary; the goal is to reduce its growth.
Our inference: Trellis should reduce owner interventions per app, including the maintenance burden of customizations. Cost optimization includes human time, model usage, recovery risk and infrastructure—not just the VM bill.
Tension/test: “you only pay for infrastructure” hides real maintenance and support costs. Track intervention minutes and recurring failures before and after adoption. Do not import Google’s staffing ratios as small-company acceptance criteria.
R14 — Entire: preserve the conversation without publishing it
Entire 0.9.0 privacy documentation and attachment implementation. Compared current documentation with installed CLI help and version-pinned behavior; a separate agent exercised synthetic capture, no-amend attachment and local-only push isolation.
Our inference: use private conversation checkpoints as supporting context for public, concise decision records. Keep remote checkpoint publication and model summaries off. Missing capture remains a stated gap, never invented history. Actual setup evidence (repository: docs/evidence/2026-09-08-provenance-setup.md) separates native-hook readiness from post-hoc capture.
The most ambitious feature-sparse product
A promise with a continuing obligation
Run proven software. Make it fit your company. Keep it working as it changes.
The distinctive bet is maintained malleability: a useful local variation should not strand a small business on a private fork. We should be judged on whether the owner can keep using, modifying, upgrading and leaving their software—not how many agents we instantiate.
This is a recommendation about where Trellis creates value. It does not replace the accepted Memos + Plane lifecycle release with a broad customization platform.
The few things a user needs
| Surface | What the user does | What Trellis must actually know |
|---|---|---|
| Apps | Install Memos or adopt their existing Plane; open either normally. | Exact deployment, supported operations, data locations, current observations, recovery and cost basis. |
| Change | Ask for a change, inspect its concrete effect, authorize it, see the result. | Versioned definition, target, permissions, relevant checks and operation history. |
| Needs you | Decide a consequential question or investigate a material exception. | Why attention is needed, observation freshness, proposed response and consequence of waiting. |
These can be views over the current modular product, not separate services. Repeated actions also need direct controls. An agent is a replaceable way to request and understand work, not the only place the system remembers it.
The developer-facing unit stays small
Keep the existing App Definition / Steward Pack. It references deployment assets, names supported operations, records data and recovery knowledge, and supplies agent guidance and behavioral tests. Do not invent a second envelope vocabulary.
For the first adaptation, keep a versioned change beside that definition: upstream version, intended difference, requested capabilities, checks, migration/removal behavior, and source/decision links. This can be ordinary files and a Git reference. Do not invent a package marketplace, policy language or graph service to distribute one modification.
Configuration first; supported extension next; a small maintained code change when needed; a replacement application only when the inherited design truly prevents the desired outcome. “Malleable” must include changing software, but it need not start by discarding it.
A concrete candidate journey
After the Memos + Plane lifecycle journey works, a small consultancy asks: “Give each client engagement the same project workflow and a linked note structure.” Trellis proposes the exact native configuration and, only where necessary, a small extension. The owner previews it and authorizes the specific change. Normal work continues in Plane and Memos.
The useful result persists without another prompt. Its definition carries the reason, upstream compatibility and checks. Another owner can import the same definition into disposable instances and supply their own configuration and authority; no client notes, IDs or credentials travel with it.
This is a candidate, not a claim that the installed Plane or Memos APIs expose all required features. Verify those extension surfaces before choosing it. If they do not, choose one smaller native adaptation; do not reverse-engineer two apps or build content synchronization just to preserve this example.
The decisive moment is the next upgrade. Trellis proves compatibility or names the exact conflict before changing the installation. A blocked upgrade with a recovery plan is sometimes correct behavior; indefinite pinning without a security plan is not maintained malleability.
The proof order
- Finish the chosen lifecycle outcome: real Memos install and existing Plane infrastructure adoption; a coherent Apps experience; a complete authorized change; monitoring; backup, restore and exit evidence. Use the active OpenSpec acceptance, not another umbrella plan.
- Prove one lasting adaptation: choose a useful supported change, apply it, use it without chat, then test upgrade compatibility and removal without deleting user content.
- Prove it travels: apply the same definition to a second disposable instance without secrets, hidden local state or the original agent transcript. Only then generalize distribution.
- Add one explicit connection when demanded: declare direction, identity, data scope, retries and deletion behavior. A shared context graph is a later projection over useful connections, not their prerequisite.
Deliberately absent
No replacement mail client in this scope, no universal company ontology, no autonomous-agent fleet, no new distributed workflow platform, no automatic production forks, no compulsory chat UI, no global marketplace, and no Rust rewrite for architectural aesthetics. Existing accepted security boundaries remain binding; the simplicity recommendation is not permission to bypass them.
What would disprove the bet?
- The owner still needs a long agent conversation for ordinary maintenance or repeated use of the customization.
- The adaptation costs more to maintain than its continuing value, or routine upgrades become private-fork rescue projects.
- A new agent cannot operate from the definition and durable state without the original conversation.
- Export exists but recovery without Trellis does not work.
- Lower hosting cost is offset by more owner toil, avoidable incidents or model spending.
Track intervention minutes, useful versus missed alerts, time to restore, compatible versus blocked upgrades, and fully attributed costs. Establish a baseline in the personal alpha before claiming numerical guarantees.
The architectural constraint that follows
We are separating three things: the application people use, the knowledge needed to operate and adapt it, and the transient agent helping with a task. Their lifetimes differ. The application and its maintained definition should outlive a conversation, a model provider and Trellis itself.
That separation is ambitious enough. Our next design choices should earn their complexity by making that promise work for two real apps and one real adaptation.