Updated

Building shared infrastructure across independent iOS products

I maintain shared infrastructure across independent iOS products, with one currently live on the App Store. This post explains how I keep that foundation coherent, what it costs, and where product-specific behavior deliberately stays separate.

I maintain several iOS products in one workspace: one live on the App Store, one larger product in development, and a set of smaller proving grounds. Each has its own Vapor backend, and all of them draw on a shared layer of Swift packages: authentication, a design system, a scheduling engine, list and shopping domains, networking. A single manifest tracks all of it: apps, backends and shared packages.

They are different products solving different problems. None of them is a variant, a localisation, or a re-skin of another. The shared layer is infrastructure and building blocks rather than a template: authentication, scheduling and a component library, not a finished app waiting for a new name.

Where that line falls is a decision I keep having to make. The shared layer stops at infrastructure: configuration, composition, feature decisions and anything that only makes sense inside one product stay in that product. When a shared abstraction starts encoding assumptions that came from one app, that is usually the signal to push the decision back out rather than to make the other apps conform to it. The foundation stays useful by staying smaller than it could be.

One of those products has shipped. Hellopost has been live on the App Store since July 2026. The others have not. The interesting claim here is not how many products exist. It is that a shared foundation underneath several products stayed coherent instead of drifting into slightly different copies of itself.

This post describes how I did that, including the parts that did not work.

The operating problem

The hard part of running several products on one foundation is not writing the code. It is that the foundation is load-bearing in more than one place at once.

A one-line change to a shared avatar component touches two independent apps, one of them shipped. A change to the deployment tooling touches every backend. An abstraction that seems obviously right for the product in front of me is often wrong for the ones I am not looking at. Every change to the shared layer carries an obligation to check those other consumers, and the cost of that check decides whether shared code is an asset or a liability.

The second problem is discontinuity. I work alone. I built the code foundation described here by hand, over years, before any coding agent was involved, and I only recently started routing execution through Claude sessions. Each session starts with no memory of the last one and ends when its task does. My own stretches at the keyboard have the same shape with a slower forgetting curve. Whatever one stretch understood has to be available at the start of the next, or the next one re-derives it. Re-deriving is slow, and it also produces different answers each time, and those differences accumulate as architectural drift.

Why chat history and ad hoc coordination failed

The obvious way to handle this is to keep the important things in my head, keep notes in a document, and keep recent context in whatever tool I was last working in. I did that for a while. It failed in three specific ways.

Context did not survive the boundary. Anything a working session established was gone by the next one. What survived was code and commit messages, which record what changed and almost never record what was rejected or why. Work would reopen a settled question because nothing on disk said it was settled.

Product decisions got settled inside tasks. When a question comes up mid-task, the cheapest move is to answer it and keep going, and most questions genuinely should be answered that way. Some of them are direction, or product scope, or a tradeoff with consequences past this task. I was resolving those implicitly inside the working session, in implementation mode, or a Claude worker was resolving them by taking the path of least resistance, instead of surfacing them as product decisions. Nothing recorded that a decision had happened, so I could not find them later.

Nothing checked the work except me. When I changed a shared package, nothing recorded whether the apps consuming it still behaved the same. The only check was my own read of the diff, and I read least carefully on the changes I felt surest about. When I said a change had not broken anything, I had my confidence and no captured evidence.

All three failures share a cause. Whoever or whatever does the work in a given stretch is temporary, and I had let truth, authority and verification live in that temporary place.

The system that emerged

The system has five parts and one rule about what survives a session. This is its shape at the time of writing, not a finished design; the parts around the worker are the ones that keep changing.

The five parts of the operating system My product-owner role sets direction and feeds two durable stores: a knowledge layer that records what is true, and a planning layer that records what is next and which decisions are waiting on me. Both hand context to a temporary worker, lately most often a Claude session, wrapped by mechanical safeguards. Everything the worker learns flows back into the two stores. DURABLE STORES PRODUCT-OWNER ROLE sets direction · rules what I reserve DURABLE KNOWLEDGE what is true about the systems every claim says how it is known describes, never plans PLANNING & DECISIONS which direction matters now which decisions wait on me what to do next THE TEMPORARY WORKER starts cold · takes one task · ends contributes, never becomes the authority MECHANICAL SAFEGUARDS Objective rules run by the tooling, so nobody has to remember them. Dashed arrows: what the worker writes back. Nothing else survives the session.
The two durable stores hand context to a disposable worker, which writes what it learns back into them.

The product-owner role is me setting direction and ruling on the decisions I keep for myself, separated out from me doing the implementation. It gets its own box because the rest of the system does not need that role present to keep working: execution continues while decisions wait for it.

Durable knowledge is a documentation repository that records what is true about the ecosystem and holds no plans at all. That restriction matters. Once a document that describes a system also holds an intention about it, a reader can no longer tell which sentences describe the current state and which describe a plan.

Planning and decisions is a separate queue with three parts: which direction matters right now, which decisions are waiting on me, and what to do next. Keeping those three apart is what lets a task name a decision as a gate instead of settling it.

The temporary worker is whatever is doing the work in a given stretch. Lately that is often a Claude session. For most of this codebase’s life it was only me, and it is still me on a Tuesday after a week away. Either way it starts with no context, takes one task, and ends. Designing for that has a side effect I care about: whatever a cold-start worker needs written down is what another engineer would also need.

Mechanical safeguards are the rules that have exactly one right answer, moved out of prose and into scripts the tooling runs, so that following them does not depend on anyone remembering them.

Underneath all five is the rule: nothing survives a working session except what it wrote down. Everything above is files in git, so it is diffable, reviewable, and portable to any tool that can read text.

Four mechanisms that carry most of the weight

The task file is the brief. Every task is one file written for someone who has never seen the conversation that produced it: the constraints, what I already ruled out, and what “done” means. Of everything here it has paid back the most.

Product decisions are a queue, not an interruption. A question I reserve for the product-owner role gets its own file. Tasks name it as a gate, and other unblocked work continues until I rule on it. The design treats my judgment as the scarcest resource: execution keeps moving while decisions accumulate, I rule on them in a batch, and a ruling recorded once can be cited later instead of relitigated.

The safeguards refuse rather than remind. A hook checks every commit to the knowledge repository against a recorded conformance baseline and rejects it if it would make the corpus less conformant. The tooling blocks a session that tries to end with uncommitted work.

Claims carry their evidentiary status. Every documented fact carries a tag saying how it is known: read from the code, asserted by me on a date, or explicitly unverified. The record does not simply say “Hellopost is live.” It says that this is my assertion, that I ran no independent deployment check, and what would have to be checked to confirm it. When a fact goes stale, the tag tells you what kind of stale it is, and therefore whether the fix is correcting a document or re-auditing a system.

What it produced

The distribution record. Hellopost has been live on the App Store since July 2026. That is the whole distribution record. RoomAid, the largest of them and the one that drove most of the architecture, is not part of that record and currently has no public users. The others are proving grounds, built to test architecture rather than to launch, and none of them has been submitted to the App Store.

One deployment path for every backend. The deployment tooling is not copied into each backend and left to drift. It is vendored from a single repository, with every project-specific value pushed out into per-project config, so the vendored files stay identical everywhere. I re-ran that check while writing this: all of them in sync. Standing up the scaffolding for a new product, its config, an empty backend, an iOS target, a database and its first migrations, is one command. What that produces is an empty project that builds, not a product: everything that makes one product different from another one is still ahead. I ran that command outside this workspace as a test, which is how I know it is not quietly wired to my own account and domain.

Products composed rather than rebuilt. The clearest case is one of the proving grounds: six screens over an event-sourced domain, drawing on the shared scheduling, design-system and calendar packages with zero changes to any of them. If composing a new product had required editing the shared packages, the shared layer would have forked in practice, whatever the repository layout said.

Verification that produces evidence. The example is small on purpose. A one-line fix to a shared avatar component, so initials shrink instead of truncating at large accessibility text sizes, reached a shipped app through the shared layer. The question for that app was not whether the fix worked but whether anything on its screen had changed. I captured its screen before and after, at the default text size and at accessibility-large, cropped the avatar region and hashed it. The hashes matched at both sizes. The rest of the screen differed in one small rectangle at the opposite end of the display from the avatar, the same rectangle at both text sizes, which is the signature of an antialiasing shift rather than of the change. The captures and the comparison are files I can open later. Before this, the same claim rested on my read of the diff.

Limits and tradeoffs

The overhead is real and it has a floor. Writing a brief for someone who was never there costs more than doing a small task directly. Below a certain size the system is not worth invoking, and I do those fixes myself.

Reuse is tested by one author. I wrote every shared package here. Every product composed from them was composed by me, or by a Claude session working from a brief I wrote. The sessions do count for something: they start with no memory of the packages, and they have built on them without my walking them through the code, so the packages are usable from what is written down rather than from what I remember. They do not show that a second engineer would be as fast. A session works from my brief and stops where the brief stops; an engineer with their own habits would push on the packages in places I have not. That is still the weakest part of the composition claim.

Verification is bounded by what I thought to check. Pixel parity proves a screen did not change. It says nothing about whether the feature was right. These mechanisms catch regressions and conformance drift. They do not catch a product decision that was wrong.

Coverage in the foundations is thin. Bringing the shared authentication and core packages under real test suites is still open work.

The agent workflow is recent and still moving. The foundation predates it by years. I started routing execution through agents recently, and I am still deciding how much of my work should run that way over the long term. The shape of the worker is the part of this most likely to change.

What I expect to keep

I am not recommending any of this as a method. One workspace and one person’s constraints shaped the specific files, and the agent side is new enough that I expect to change it. The parts I expect to keep, whatever ends up doing the executing, are the ones that do not depend on an agent at all: the three questions kept in three places, briefs written for someone who was not in the room, rules with one right answer moved into tooling that refuses, and verification that leaves a file behind.

What I want from all of this is narrow: a codebase that stays coherent while several initiatives run through it, and that does not depend on my memory to stay that way.

If you are working on something like this

Email me and tell me what you are working on.

← All posts