← Back to blog

How PMs and Designers Started Shipping Our Mobile App

When we rewrote our mobile app from Flutter to Kotlin Multiplatform, we set out to build a faster, more stable app. We ended up with something we hadn’t planned for. By the end, the people shipping features weren’t only engineers. Alongside our six engineers, two product managers and two designers were opening pull requests and merging real features on their own. None of them had to become engineers to do it. The tooling we built around the AI took on enough of the coding that their part was deciding what to build and confirming the result was right.

Our architectural shape

The codebase has a consistent shape. A shared data layer sits under the whole app, and each feature is split into two modules, an api module that declares what the feature offers and an impl module that holds how it works. One feature depends on another through its api and never reaches into its internals.

Inside a feature, the work is split into small pieces that each do one job. Two UIs sit at the top, one written in Compose for Android and one in SwiftUI for iOS, both drawing from the same layers below. A renderer maps screen state into something those UIs can draw, and a presenter produces that state. Under them, a use case holds the business logic, and a repository decides what data to hand up, pulling from a local cache and the shared data layer.

The shape of a feature, repeated across the app An outer box labeled App holds everything. On the left, a box labeled shape of a feature expands one feature into its layers, two UIs in Compose and SwiftUI over a renderer, a presenter, a use case, and a repository with a local cache, above a shared data layer. A vertical divider separates it from a zoomed-out column of small boxes labeled feature A, feature B, feature C and more, all inside the same App box. Outside the app, past another vertical line, three more boxes labeled design system, useful repositories, and Figma config. APP SHAPE OF A FEATURE Compose UI SwiftUI Renderer Presenter Use case Repository Local cache Data layer VERY MUCH EVERY FEATURE Feature A Feature B Feature C Feature D Design system Useful repositories Figma config
One feature expanded on the left, the same shape repeated across the app. Beside it sit the resources the tooling draws on, the design system, useful repositories, and the Figma config.

This is more or less the shape of every feature in the app. Zoom out, and the codebase is that one pattern repeated, one feature after another, each cut from the same layers. That regularity is what lets a skill encode a feature once and someone new pick it up without holding the whole system in their head.

Enforcing patterns

From the diagram above, we can see that the code structure of each feature looks the same. Most of a new one is boilerplate, the same files wired the same way. That part is deterministic, so a set of scripts and a CLI generate it with no AI involved. Creating a skeleton does not need a model, it needs a template.

The AI only comes in for the parts that are not deterministic. Every layer has a skill that spells out what to do and what to avoid, how to find the right API to consume, and when to cache a result and when not to. The AI is wired to call the CLI for the deterministic parts, so it works on those choices and leaves the mechanical files to the scripts. Together, the scripts, the CLI, the skills, and the rules make up what we call the harness, a layer we built on top of the model during the rewrite to generate code the way our project expects.

To push the accuracy further, the design system has its own rules, and one of them tells the AI to go read the design system’s codebase when it is unsure instead of assuming a pattern. Reaching for the source on a doubt keeps it from reinventing a component that already exists.

None of this is fixed. We revisit the skills, commands, and rules as the models improve, and a newer model tends to need fewer of them than an older one, since it gets more right on its own. The set shrinks as the models get better at holding the patterns without being told, and grows as the codebase raises new needs.

Figma MCP

Figma plays an important part here. With the model instrumented the way described above, a Figma asset carries a large share of the context a screen needs. It holds the shape of the screen along with its text and what each part is about, which gives the AI a much clearer picture of what to build.

Figma Code Connect, paired with a design system whose components map one to one to their Figma counterparts, tied each design straight to real code and enriched what the AI produced. An HTML mockup gives none of that. Its markup maps to nothing in the design system, so the AI would sometimes rebuild a component that already existed. None of this is locked to Figma. Another tool like Paper.design would fill the same role, and the MCP itself could be swapped for a CLI to cut token use, though with Code Connect it was already efficient.

Prompting a feature

With the shape, the skills, and the Figma connection in place, building a feature comes down to prompting one. A contributor opens an AI coding agent in the terminal, points it at the Figma design for the screen, and describes what the screen should do. From there the harness takes over, calling the CLI to scaffold the deterministic structure and leaning on the per-layer skills for the parts that are not.

In practice a prompt could be this loose.

Let’s build a new screen [figma link], this is the chat history. It has an entry point on the chat screen, and the chat history should keep a local cache.

That was usually enough for a good result, in part because the data layer for most domains was already in place, even for endpoints we weren’t calling yet, so a new screen usually found its data waiting. When something was missing, the AI would ask for the context it still needed, a cache strategy or the OpenAPI spec for an endpoint, before it started building.

The person driving this wasn’t hand-writing Kotlin. The AI was. Their part was deciding what to build and whether the result was right, without needing to know how the pieces fit together underneath. And when they got stuck, on a review comment, a failing check, a bit of wiring that made no sense, the fix was another pass with the AI rather than a crash course in the framework.

Onboarding

Once the tooling was in place, we invited a few people to try working in the codebase. We looked for the ones with a curious streak, reached out to them directly, and onboarded them ourselves. For any of that to work, setup had to be simple.

Configuring the project is a single command. Running ./gradlew setupProject does the whole job and leaves it ready to build.

Getting the machine itself ready is the other half, the tools and accounts a mobile project needs installed and set up. For that we shipped a set of install scripts and a prompt that has the AI walk a new contributor through it, step by step, until everything is in place and the app builds. No engineer required.

Trying it out and merging

Validating what the AI generated meant getting the app onto a device, and for a non-engineer that usually meant a full local build. On a codebase this size, and depending on the machine, that could leave it crawling for a good while.

Shopify’s Tophat took that pain away. It is a neat tool that runs an in-development build on an Android or iOS emulator, or on a real device, and installs it in one click. Once someone has Tophat set up on their machine, checking whether a feature looks right comes down to clicking a link, which installs the app for them. The repository has the details.

What some of the non-engineers did with it caught us off guard. To avoid bogging their machine down, they leaned on a /tophat command we had built into the project. Dropping it in a pull request comment kicked off a CI build and handed back an installable version to check, with no local build at all. Running locally is faster, but since shipping code was not their main job, letting CI build while they carried on with everything else suited them well.

From there the change went through the same gates as any other. A quality gate in CI ran the checks, an automated review pass flagged what looked off, and the contributor refined the code against that feedback before it merged. It was the same loop an engineer would run, with the harness and the review carrying the parts a non-engineer wouldn’t know to check.

From a design to a merged feature A mockup or Figma design feeds the AI and the harness, which generate the feature. It then goes through a test build, a quality gate, and an automated AI review. Refinement loops back to regeneration before the change merges. Mockup or Figma AI + harness Test build Quality gate AI review Refinement repeat
A design goes in, the harness generates the feature, and it runs the same test build, quality gate, and review loop as any other change.

The engineer’s role

Engineers did not go away. What changed was the role. Day to day, one still got pulled in for two kinds of things.

The first was vocabulary. Our UI is built from a system of widgets, and a contributor would naturally describe what they wanted in plain terms, like asking for a card on a screen. Knowing that the card is really a widget in our system, with an API already built for it, was context a developer could hand over in a sentence. That one sentence kept the AI from building something new where a block already existed.

The second was polish. Every so often a feature shipped needing a little more finish, or with a small piece missing. It was the exception rather than the rule, and usually a quick follow-up, but it is why the review step earned its place.

The bigger job sits underneath all of it. Engineers own the structure and the tooling, the scripts, the skills, the rules, and the boundaries that keep every feature consistent and the whole thing able to scale, and they keep it current as the models and the codebase move on. They stay the gatekeepers of quality, and building that tooling is what buys the speed. With the rails solid and maintained, a PM, an EM, or a designer can actually ship code that makes the product better and faster.