A reasonable thing to ask before trusting any “AI that learns your business” claim: learns it how, exactly, and how would anyone know if it actually worked? Most of the time the honest answer is that it doesn’t. The model is the same model on day one and day three hundred, and any sense of improvement is really just you getting better at prompting it.
We wanted a real answer for our own product, not a comfortable one. So we tested it two different ways.
The slow way: a simulated company, twelve months
AlphaForge is a brand studio we use as a simulation: a realistic twelve-month arc run against the live product, with real documents and real usage patterns, even though it isn’t a paying customer. At the start, the same question, “what should we charge for a project like this,” got a generic answer, because there was nothing in the Vault yet to ground it in.
Six weeks in, whenever the founder actually sat down to do it, not on a fixed schedule, came the first structured Monthly Alignment. By then the Vault wasn’t actually starting from nothing: raw notes and the activity already gathering from real client work had been accumulating quietly the whole time. The Alignment gave that a shape. The next time the pricing question came up, the answer was different: it referenced an actual documented pricing philosophy, mentioned a specific recent client by name, pulled in a methodology written down once, in a fifty-word raw note, weeks earlier. By month six, it was citing patterns across four past projects and a specific team observation about how a particular kind of client behaves. By month twelve, drafting an offer that used to take six hours took forty-five minutes, not because anyone got faster at typing, but because the system already knew the scope, the objections, and the wording that tends to work.
That’s the mechanism, plainly: every Monthly Alignment, every team observation, every offer pattern, and the ordinary daily activity in between, gets written into the Vault, and every future answer gets assembled with that history as grounding context. More accumulating doesn’t mean noisier answers, either. Curatoria, the Vault’s own maintenance agent, curates and summarizes as it grows, archiving rather than deleting, always with a human’s approval first. It’s not a bigger model. It’s the same model with somewhere real, and increasingly well-kept, to look things up.
The fast way: a controlled month, checked against a real bill
The twelve-month story is convincing but slow to verify independently. So we also ran something faster and colder: a fully controlled, one-month simulation. Four people, one founder, three team members, real inquiries, real client conversations, real offers, every single AI interaction a genuine call to a live model, not a script reciting pre-written answers.
Two things came out of it.
The cost was negligible. We checked the AI bill before and after. The entire simulated month, four-person team, cost under a dollar, total, not per user. The fear that AI features quietly become an expensive second subscription doesn’t hold up at this scale, even generously multiplied for a much busier team.
The honesty held up under pressure, which is the part we actually care about. At the end of that month, we asked the system’s strategic review layer to look back at what had happened. Seven offers had been drafted and approved, a month that looked busy on paper. A tool built to keep you engaged would have called that a great month. Ours didn’t:
“Six months in, the company is still only writing proposals… Pipeline illusion: high volume of approved drafts creates a false progress signal while zero revenue or client work materialises.”
Nobody asked it to find a problem. It looked at the actual pattern, drafts going up, nothing converting to paid work, and said so, plainly, in language that would make anyone slightly uncomfortable reading it about their own month. That discomfort is the point. A system that only ever tells you good news isn’t actually paying attention.
Why this matters more than the headline number
It would be easy to lead with “our AI gets smarter over time” and leave it there, because that’s the comfortable version. The less comfortable, more useful version is: a memory system is only worth trusting if it’s also willing to be the one telling you something you didn’t want to hear. We built the compounding specifically so the system would have enough real history to notice patterns like “this keeps not converting,” not just to sound more impressive in a sales call.
The bet underneath the compounding
There’s a bigger question sitting under all of this, worth saying plainly: what happens when every business starts running on AI trained to give everyone roughly the same answer? Generic-by-default AI doesn’t just risk being wrong sometimes. It risks making every business that leans on it start to look the same, sound the same, make the same calls, because the model has no memory of what actually makes this company different from the one next door.
The bet underneath Craft11’s compounding runs the other way. Not a bigger model giving smarter generic answers, but the same model with somewhere real, specific, and private to look things up, learning with a business instead of just about business in general, cheaply enough that running it is never a tradeoff. Under a dollar for a month of real usage isn’t just a nice number. It’s what makes it possible to give a small team the kind of structured institutional memory only large companies used to afford, without taxing them for the privilege.
That’s also why the AI here was never meant to be the endpoint. It’s the starting one. Picture fifty different companies running Craft11 across fifty different industries, three or five years from now. The interface will look almost identical on all fifty screens. What’s actually inside each one, the pricing philosophy, the client patterns, the hundreds of small judgment calls accumulated one Monthly Alignment at a time, will have nothing in common. That’s the point, not a side effect. AI here is meant to support the diversity of how fifty different leaders choose to run their businesses, not flatten it into one “best practice” everyone gets nudged toward. The compounding is the product. The AI is just what makes the compounding possible.
Twelve months of AlphaForge’s story shows what compounding feels like from the inside. One real month of checked-against-a-bill simulation shows what it costs and whether it stays honest under pressure. We’d rather show you both than ask you to take either on faith.