The LLM Was The Easy Part
Building a local AI app taught me the hard part was everything around the model.
I thought I was building a local chatbot.
Private conversations. Legal research. Sensitive notes. Things I did not want sitting on someone else's server. The plan sounded simple enough: run a model locally, save the chats locally, make it remember me.
I expected model selection and local inference to consume most of the project.
Then I installed Ollama, pulled a model, sent it a message, and got a response.
Ollama gave me a chat-shaped API. RecallMEM still needed to save, reload, search, delete, and recover the data around each response.
If an AI app only has to demo, you can fake a lot. If you plan to use it tomorrow, normal software engineering shows up immediately. Saving conversations. Loading them without corrupting them. Remembering facts without inventing them. Uploading PDFs. Rendering markdown. Deleting data for real. Letting someone install it on a machine that is not yours.
The model call was boring
Ollama made the first model call boring in the best way. Send messages to a local HTTP endpoint. Stream tokens back. Put those tokens in a chat bubble.
There were details, but they were contained. Gemma had thinking mode enabled by default, so it spent tokens writing out a reasoning block before answering simple prompts. Adding think: false made normal chat much faster. The 31B model was too slow for my taste, so I switched to a 26B mixture-of-experts model and used a smaller model for background tasks like title generation and fact extraction.
Those tuning problems stayed contained.
The chat itself needed messages, titles, timestamps, model metadata, provider selection, attached files, partial saves while streaming, a stop button, pinned chats, renamed chats, and recovery when the browser closed at the worst possible time.
Deletion was harder than building
My first instinct was to reuse Speak2Me, the production voice AI journal app I had already built. It had memory, prompts, facts, transcripts, embeddings, and a real product shape. I figured I could strip out the cloud parts and keep the good stuff.
That was wrong.
I started deleting Hume, Stripe, auth, Inngest, voice components, and production-only routes. Every file I deleted broke five files that imported it. The graph page depended on the journal page, which depended on the dashboard, which depended on the background pipeline. Within an hour the codebase looked like a half-disassembled engine.
Speak2Me's code had optimized around the product it became. Removing one feature exposed dependencies across the rest of the application.
I stopped fighting it and started fresh with a new Next.js app. The only thing I kept was a reference folder with the parts that were actually valuable: prompts, memory extraction logic, types, dates, and fact handling. Not imported. Just reference material.
The fresh app let me reuse the memory logic without carrying every dependency from Speak2Me.
The database fight was the warning
The next fight had nothing to do with language models.
I needed storage. SQLite was tempting because one file and zero setup is hard to beat. TiDB matched Speak2Me, but self-hosting a distributed database for a personal chatbot made no sense. YugabyteDB was interesting because it speaks Postgres and supports pgvector, but the Homebrew tap path was dead and the install options were more work than the app deserved.
I chose vanilla Postgres with pgvector.
That choice kept the future path open. Local Postgres today. Maybe managed Postgres later. Maybe distributed Postgres later. Same driver. Same SQL shape. Same vector operators.
Then Homebrew made it annoying. Postgres 16 did not have the pgvector files I needed. Postgres 17 had stale share directories and mismatched libraries left behind from prior installs. I ended up wiping the broken Postgres pieces and reinstalling cleanly before CREATE EXTENSION vector; finally worked.
Until pgvector loaded correctly, RecallMEM could not store or retrieve the vectors used by its memory system.
Local models still have knobs
Local AI sounds like one decision: run the model on your machine.
In practice, it is a pile of smaller decisions. Which binary is actually running? Which server is the CLI talking to? Is the desktop Ollama app running one version while Homebrew installed another? Is thinking mode on? Is the model dense or mixture-of-experts? How much prompt are you sending every turn?
I had Ollama installed twice. The server and client versions did not match. The symptom was not a clear error. A model pull just failed with a link to the download page buried in output. Once the versions matched, the same command worked.
Running locally gave me privacy and control, along with responsibility for every broken dependency on the machine.
The installer now had to detect the environment, explain what was missing, recover from bad defaults, and stop making users debug assumptions from my laptop.
Memory made it a system
The first memory bug looked harmless until I checked the database.
I saved conversations as text transcripts. Each message started with user: or assistant:, and messages were separated by blank lines. That looked readable and simple.
Then markdown happened.
Assistant replies have blank lines inside them. Headings, lists, paragraphs, code explanations. My loader split the transcript on blank lines, kept blocks that started with assistant:, and dropped continuation blocks without a prefix. The full response was in Postgres, but loading the chat silently threw away most of it.
Worse, if the user kept chatting after a reload, the truncated conversation could get saved again. A parser bug became memory loss.
I changed the loader so continuation blocks attach to the previous message. Before that fix, save and load disagreed about the transcript format and silently discarded content.
Background extraction failed a different way. I clicked New Chat two seconds after a response, and the next chat loaded before the facts were saved. RecallMEM now flushes that work when the user changes chats.
The Interface Started Losing Work
I treated chat UI details as polish until long generations, forced scrolling, and lost drafts interrupted real work.
A local model can think for a minute, so the stop button matters. Streaming can pull the page away from what I am reading, so sticky scroll matters. A refresh can erase a draft, so recovery matters.
PDF upload expanded from text extraction into worker bundling, scanned-page vision, and diagram handling after each simpler version failed.
The model did not care about any of this. The user did.
Installability is product
RecallMEM worked on my machine for the least interesting reason: my machine had months of development state on it.
Postgres was installed. pgvector worked. Ollama was running. Models were already pulled. Environment variables existed. Weird one-time setup problems had been solved and forgotten.
Then I tested the npm package on a clean laptop and hit three showstoppers in 30 minutes. Ollama was installed but not running. The model picker was skipped because of a setup flag. Background memory extraction used a hardcoded fast model that the laptop did not have.
The app was not installable. It was lucky.
I added npx recallmem to detect Postgres, pgvector, Ollama, the database, services, models, migrations, and environment files. To keep the npm package small, the CLI ships as a bootstrapper and clones the application during setup.
The published package was about 22KB because it only bootstraps the real application. On a clean machine, setup became the first product screen whether I designed it that way or not.
What I learned
The first model response arrived before any of the failures in this post.
RecallMEM still had to remember correctly, delete honestly, recover from reloads, explain setup failures, control token costs, and install on a machine I had never touched.
The LLM was the easy part. The product was everything I had to make reliable around it.
Questions about this post? Ask the terminal on my homepage — it knows this whole site.