By Yaokai Jiang · Founder & CEO, Momen Talk delivered at No-Code Week Frankfurt · June 16, 2026
On June 16, I gave a talk at No-Code Week Frankfurt about something our team has been wrestling with for a while: what it actually takes to bring AI agents into a no-code platform — not just for a demo, but all the way to the finish line. The title was Vibe No-coding: Challenges in the Pursuit, and it was an attempt to be honest about the tradeoffs, the dead ends, and the approach we've landed on. What follows is that talk, in my own words.
The dream is simple: describe the app, watch it appear. And prompt-to-app does deliver on that, fast. Great time-to-value, easy mental model, great for exploration and demos. The problem is there's no finish line.

We've written about this pattern before: why vibe coding breaks at scale, and why one prompt can't build your startup. The short version is this: when you use AI to generate against our platform, the fundamental challenge is that the AI doesn't know — and can't intuit — our DSL. If you just ask the AI to directly write our DSL, it's going to make a bunch of mistakes, because there are so many inherent assumptions embedded in it. It's not just logic like most code — there are platform assumptions, one-time behaviors, all of that already baked in, and that's very difficult to convey to the AI. So we can't do the spray-and-pray method. You can't just fire off agents generating a bunch of code that's probably going to fail.
And it's not just about putting too much in the prompt and suffering from context. Even if you use something like Claude Opus 4.8, which actually works pretty well even at around 5K context, you still have to balance three things that are always going to be there when you use AI: accuracy, cost, and latency. If you have 500K context, it's going to cost a lot — $2.50 every time you invoke it. And assuming you consume 1,000 tokens a second at the ingestion stage, that's 500 seconds before you have the first token back. You don't want any of that. A giant prompt is simply not going to work.

Momen's visual builder is a UI over a structured language — our type system ztype, shared front-to-back across three pillars: Design, Data, and Action. The diagram is the program, not generated code. This is what makes it comprehensible to non-technical builders.
But it's also what makes agentic adoption hard. The model was never trained on our product language. The DSL carries hidden assumptions — runtime behavior, valid combinations, migration rules. Spray-and-pray DSL generation looks plausible, yet violates type, dependency, and runtime constraints. This is the core tension we've had to engineer our way through. If you're thinking about how AI coding and no-code compare for non-technical founders, this distinction is exactly where those two worlds diverge.

So what do we do? I came up with something — I'll admit it's a bit unconventional. It's essentially a plugin system that runs on a local server, borrowing heavily from how we structure our own codebase.
The system has a God prompt plus a set of plugins. Each plugin contains three things: knowledge to work with a particular subsystem, a set of tools we expose to the agent to modify our DSL, and a set of checks that run before and after the AI decides to perform a tool call.
The God prompt — we also call it the secret loop — is responsible for deciding which part of the system we're currently working on, and ensuring the right bucket is loaded. That means the agent has domain knowledge plus the tools to search and operate on that particular facet of the system.
These are the actuators we've built and exposed as plugins: database, type, binding, log, component, actionflow, permission, payment, docs. A tool cannot even be called until its area is loaded.

Once the right plugin is loaded and a tool call is issued, what comes next is immediate feedback — and that's critical for good agent performance. Just like us: we learn to ride a bicycle fast because feedback is immediate and deterministic. Delayed, slow, random feedback doesn't train well. In our loop, deterministic feedback lets the AI correct its course right away if it makes a mistake.
I'm still working out exactly what to feed back to the AI, and how to balance that with context efficiency. There's always a cost associated with the amount of feedback. For now, I'm sending everything back, but I'll share some tricks at the end.
The most important part of this approach is making sure that the AI stays at the same abstraction level as the human it's supposed to serve. If we're serving technical builders, use code and review code. If we're serving non-technical founders, don't produce code — produce visual artifacts they can actually read. Every single step the AI performs should be comprehensible to the human reviewing it. It becomes auditable. The sequence makes sense to the human. And the human actually learns along the way — otherwise, the information just glazes over and nothing sticks.
I want to be upfront: this is not a perfect solution. It is going to be a slower loop than just generating code. Instead of saying "write that entire class" — one output, single cost — you now have multiple rounds, multiple function calls, more cost involved.
And every time a new model is released, the performance uplift is not as immediate for us. New models have been trained on more code, optimized for better coding, and while some of that architectural understanding might generalize to a no-code environment, what ultimately drives performance here is the parameter and assembly work — how we construct the JSON, what knowledge we give it, what constraints we impose.
But we took those tradeoffs deliberately, not out of desperation. Our view is that we're here to help non-technical founders build and scale their products — not to stop three steps before the finish line, but to actually cross it. Reliability is more important to us than flashy demos. Prompt to 70% is easy. Getting to 100% is the hard part. This is the same reason we believe building fast with AI doesn't automatically mean you can launch fast — the last mile is engineering, not generation.

If you're building anything that involves AI agents, here's what I've learned about context management:
Load on demand. Don't have one giant dot-prompt. This is essentially progressive disclosure — any time you need a new piece of information, have the AI retrieve it on the spot. The God prompt should only carry the overarching context; specific domain knowledge gets activated per subsystem.
Bound your outputs. Otherwise you'll end up flooding your context with a single massive tool response. We cap logs at 100 rows, sub-loops at 10 turns.
Don't truncate. Compact instead — summarize old turns, keep goals and schema names verbatim.
Externalize state. Don't rely solely on message history to reconstruct the current state. We externalize three things: user-level preferences, project-level information (what does the project do, user stories, a to-do list), and current verification status.
Keep the cache prefix stable. This one is particularly neat. The largest models have automatic prompt caching — if the prefix of your message is identical, it gets cached. The stuff you want to manipulate is volatile, so put it at the tail. Everything before that stays cached. Keep the volatile nudges — to-dos, current status, verification results — riding at the end. That alone will save you around 90% of cost and make things roughly 90% faster.
The shape of the message history we've settled on looks like this:

Role | Content | Cache behavior |
|---|---|---|
system | role · memory · platform overview | cached |
summary | older turns, compacted | cached |
user | "add urgent-ticket escalation" | appended |
assistant |
| appended |
tool | knowledge + tool schemas | appended |
assistant |
| appended |
tool | diff · dry-run · logs | appended |
nudge | to-dos · status · verify | appended |
The stable prefix stays cached. The agent only appends to the tail.
We're not building toward a demo. We're building toward something a non-technical founder can actually own, debug, explain, extend, and trust. That's the real challenge in vibe no-coding — not interpreting intent, but turning intent into constrained, visible, testable product operations. For a broader look at how AI builders compare for real data and workflows, we've covered that on the Momen blog as well.
Come find me at momen.app, or connect with me on LinkedIn. Let's cross the finish line.