Building Once Upon Today: Why I Tested My AI Pipeline Locally First

A Quick Update on Once Upon Today Before We Get Into the Build

Before getting into the local AI pipeline, here is a quick update on where Once Upon Today currently stands.

I recently redesigned the app’s logo. The original version was AI-generated, which helped me establish a direction quickly, but it had one problem I could not unsee: the bottom portion was supposed to resemble an open diary, yet it looked more like a leaf. It had two curved lobes but no clear spine or pages. It was a pleasant shape, just not the right one.

Old logo. Old logo.

So I sat down and redrew it by hand — my first time doing vector design in Affinity Designer, after years in Adobe Illustrator. The shortcuts and tools took some relearning, but one thing made it easier than expected: Affinity ships with ready-made shapes for crescents and stars, which happen to be exactly what’s already in this logo’s night-sky design. That saved real time.

The result: the same core idea — a crescent moon, a scatter of stars, an open book, a quill — but the book now actually reads as a book, and the whole thing feels more intentional than the AI-generated first pass.

New logo. New logo.

Small update, but it’s the kind of detail that matters when this logo is going on an app icon, a splash screen, and a Play Store listing. Building Once Upon Today entirely in public for Shipaton 2026 — more updates as the last stretch to the deadline continues.

With the visual identity taking shape, the next challenge was building and repeatedly testing the system at the heart of the app: the three-stage AI pipeline that turns a voice note into an illustrated diary panel.

The setup

Transcription runs on whisper.cpp, installed via Homebrew, with a small base.en model pulled directly from Hugging Face and served locally with whisper-server — a plain local HTTP endpoint I can curl an audio file at and get a transcript back.

The prompt-writing layer runs on Ollama, serving Qwen2.5 (3B). This is the piece that’s easy to overlook: a transcribed voice note (“ugh today was rough, barely slept, then the meeting got moved again”) isn’t an image prompt. Something has to rewrite it into a short, concrete visual scene description before it ever reaches the image model. Ollama made this almost free to stand up — it’s already a local HTTP API out of the box, so there was no custom server to write, just ollama pull and ollama serve.

Image generation was the one that actually took two tries. I started with AUTOMATIC1111/Stable Diffusion, and hit a wall of unmaintained-repo problems — a dead submodule, pkg_resources/setuptools incompatibilities, Python version pinning fights. None of that is a technical dead end, just a time sink I didn’t have. I switched to FLUX.2 [klein] (4B), served through a small diffusers-based FastAPI app: Apache-2.0 licensed, actively maintained, small enough to run reasonably on a laptop, and no legacy build issues to fight. First run pulls the weights from Hugging Face automatically — no manual checkpoint dance.

All three are wired into a single startup script, triggered with one command — so a dev session starts with one line instead of three separate terminals.

Why local, not a real provider

What I was actually rebuilding wasn’t the AI — it was the orchestration around it: batching moments, reserving and refunding credits, handling a failed generation without losing an entry. None of that logic cares whether the image came from Imagen or a local model, but it needed testing dozens of times a day. Doing that against paid APIs means paying roughly in proportion to my mistakes, not my usage. Rate limits and network round-trips add up fast during rapid iteration too, and it kept diary-like test content entirely on-device during development.

The honest caveat: local output doesn’t match production quality, so this validates the pipeline, not the final visual bar — I still test against the real Google APIs periodically, since illustration quality is core to the product itself.

At one point, all three of these were running at once alongside live coding and video editing on the same machine — the actual story behind the laptop-stress-test Reel, not chaos for its own sake.

Once Upon Today is being built entirely in public for Shipaton 2026, RevenueCat’s global hackathon. The waiting list is open.

Leave a Comment