Today's AI models can do almost anything. Content strategy. Legal case prep. Financial audits. Code reviews. UI engineering. Cybersecurity analysis. Compliance documentation. The list keeps growing every quarter, and at this point the ceiling of what these models are technically capable of is far above what most people ever ask them to do. But here is the thing nobody wants to talk about. The people getting the most out of AI are the people who needed it least.
Domain experts. The lawyer who already knows which clauses a contract is missing, so she tells the AI exactly which ones to check. The senior engineer who already knows about idempotency keys and race conditions, so he asks the model to handle them. The auditor who has seen a thousand balance sheets and knows exactly which line items to flag. These people are getting dramatically better results from the exact same tools everyone else has access to. Not because the AI is smarter for them. Because they are asking smarter questions.
The real gap
The AI already knows everything a domain expert knows. It has been trained on it. But it will not use that knowledge until it is explicitly asked. That is the difference between an expert using AI and getting extraordinary output, and a regular user typing a vague request and getting something generic back. A 2025 study from MIT found that nearly half of the performance gains users see when switching to a better AI model come not from the model itself, but from users adapting how they write their prompts (MIT Sloan, 2025). Half the improvement is the human, not the machine.
This is not a new observation. In 2023, researchers from Harvard, Wharton, and MIT ran a controlled experiment with 758 BCG consultants using GPT-4. They called the result the "jagged technological frontier." For tasks inside AI's capabilities, consultants using AI saw output quality rated over 40% higher. And the biggest gains went to lower-skilled participants, who saw a 43% performance increase compared to 17% for top performers (Dell'Acqua et al., Harvard Business School, 2023). The takeaway is clear: AI is a leveler, but only if you know how to use it. And most people do not.
The UK government's own survey, conducted by Ipsos in 2024, found that only 28% of the public feel confident in their ability to use AI tools in their daily life, despite 73% having used AI in the previous month (UK DSIT/Ipsos, 2024). That is a massive population of people using tools they do not feel equipped to use well. So what happens? They type something vague. Something ambitious but underspecified. The AI, being a generalist by default, assumes half the context. It fills in the blanks with safe, generic defaults. And it gives them something that is close to what they wanted but nowhere near what the model was actually capable of producing. Then they blame the AI.
The bottleneck is the user, not the tool
I am not saying this to be harsh. I am saying it because it matters. The actual bottleneck in most AI interactions today is the human on the other side. Not the model. Not the tool. The person using it. McKinsey's 2025 State of AI survey found that 79% of organizations are now regularly using generative AI, but only about a third have managed to scale it to real enterprise value (McKinsey, 2025). The gap between adoption and actual value extraction is enormous.
The 2025 Stack Overflow Developer Survey tells a similar story from the practitioner side. 84% of developers now use or plan to use AI tools. Over half use them daily. But only 29% trust the accuracy of the output. And 66% report spending more time than expected debugging or fixing "almost-right" AI-generated code (Stack Overflow Developer Survey, 2025). People are using these tools constantly and trusting them less every month. Something is broken, and it is not the models.
The industry's smartest people are starting to name this problem. In June 2025, Andrej Karpathy, one of the founding members of OpenAI and former head of AI at Tesla, wrote that the real skill is not prompt engineering but "context engineering," which he defined as "the delicate art and science of filling the context window with just the right information for the next step" (Karpathy, X, June 25, 2025). Shopify CEO Tobi Lütke framed it as "the art of providing all the context for the task to be plausibly solvable by the LLM" (Lütke, X, June 19, 2025). Ethan Mollick at Wharton, whose research on AI and work has shaped how entire industries think about this, puts it even more directly: the most effective approach is to treat prompts like management tasks. Clearly specify the goal, define what good and bad output look like, and establish ways to test the results (Mollick, Wharton Generative AI Lab, 2026).
By 2026, Gartner is positioning context engineering as a major enterprise discipline, predicting that 40% of enterprise applications will integrate task-specific AI agents by year-end (Gartner, 2026). But they are also saying something important: most AI agent failures in production are not caused by poor prompting. They are caused by context failures, where the model lacked necessary information or was drowned in irrelevant data. All of these people are describing the same problem from different angles. The models are powerful enough. The bottleneck is the space between the user and the model.
What Kosmo actually does about this
This is the problem that nobody else seems to want to focus on. Everyone is building better models, building better tools, building better wrappers. But the person sitting in front of the screen, trying to get a good output, is still stuck figuring it out on their own. And the problem compounds. Every model update changes what works. Every new tool ships with its own conventions. The documentation grows faster than anyone with a real job can read it. The expert stays current because keeping up with AI is part of their work. The regular user falls further behind with every release cycle, not because they are not smart enough, but because they have their own job to do. That is not a skill gap. That is a structural problem. And telling people to "just learn prompt engineering" is not a solution. It is passing the burden to the person least equipped to carry it.
I built Kosmo to fix exactly this problem. Not by building another AI wrapper. Not by giving users more templates to copy. By doing the work that the expert would have done before the user ever hits send. Kosmo is a prompt compiler. You give it what you want in plain language, the way you would explain it to a smart coworker. Kosmo's agents (five of them, working together) take that raw intent and compile it into the kind of structured, tool-aware, domain-enriched prompt that a power user or domain expert would have written. The kind that pulls far more of the capability out of whatever AI tool you are sending it to than a vague prompt ever would.
It is not an API wrapper. It is a foundational system. Multiple agents analyzing your intent, enriching it with domain knowledge, structuring it for the specific model you are targeting, and catching the things you did not know to ask for. Every single user gets this, regardless of their experience level. The goal is simple: reduce the difference between the output quality an expert gets and the output quality a regular user gets. Make a normal user a power user. Not by teaching them prompt engineering. By doing it for them.
Where we are right now
I will be honest about where we stand. We made the first version of Kosmo in 69 days with almost no resources. We went live for public beta on August 6th. One week later we had 150+ users, all of them manually connected with, all of them giving us direct feedback. Zero marketing spend. As we test it ourselves and collect feedback from early users, we are finding that outputs created using Kosmo almost always beat outputs without it. Does it produce great output 100% of the time? No, and we do not claim that. But the signal from our early users is clear and consistent: the compiled version is better than what they would have written on their own.
We have not posted benchmarks yet. That is intentional. We are still tuning based on real user feedback, not trying to game a leaderboard. When we do benchmark, it will tell the actual truth about where we stand. And by then, we plan to be significantly better than we are today. The next version of Kosmo will be built on everything we are learning from these first 150+ users. And as we grow and get more resources and more intelligent people involved, it is going to become something this AI era needed from the beginning. A layer that makes every AI tool work the way it was supposed to. Not just for experts. For everyone.
Sources:
- MIT Sloan (2025). "Study: Generative AI results depend on user prompts as much as models." mitsloan.mit.edu
- Dell'Acqua, F., McFowland, E., Mollick, E. et al. (2023). "Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of AI on Knowledge Worker Productivity and Quality." Harvard Business School. ssrn.com
- UK Department for Science, Innovation and Technology / Ipsos (2024). "AI Skills for Life and Work: General Public Survey Findings." gov.uk
- McKinsey & Company (2025). "The State of AI in 2025: Agents, innovation, and transformation." mckinsey.com
- Stack Overflow (2025). "2025 Developer Survey." stackoverflow.co
- Karpathy, A. (June 25, 2025). Post on X defining "context engineering." x.com
- Lütke, T. (June 19, 2025). Post on X on "context engineering." x.com
- Gartner (2026). "Lead the Shift to Context Engineering as Prompt Engineering Fades." gartner.com
Frequently asked questions
Why do experts get better results from AI than beginners?
Because experts already know what to ask for. They have the domain knowledge to tell the AI exactly what to include, what to check, and what to avoid. The AI has that same knowledge in its training data, but it will not surface it unless the prompt specifically asks for it. A vague input gets a generic output. A precise input gets a precise output. The difference is the person, not the model.
What is the difference between prompt engineering and context engineering?
Prompt engineering is about crafting clever instructions for a single interaction. Context engineering, a term popularized by Andrej Karpathy and Shopify CEO Tobi Lütke in 2025, is broader. It is about managing everything the AI needs to know before it generates a response: the right background, the right constraints, the right structure, the right format for the specific tool you are using. Most AI failures in production are context failures, not prompting failures.
Do I need to learn prompt engineering to get good AI output?
That is the current expectation, and it is the wrong one. Prompt conventions change with every model update. Each AI tool has its own format preferences. Keeping up with all of it is a job in itself. The real question is whether users should have to carry that burden at all, or whether a system can handle it for them. That is the problem Kosmo was built to solve.
What is Kosmo?
Kosmo is a prompt compiler. You describe what you want in plain language. Kosmo's five agents analyze your intent, enrich it with domain knowledge, structure it for the specific AI tool you are targeting, and catch what you did not think to ask for. The compiled prompt then goes to whatever AI tool you choose. It is not a chatbot. It is not a wrapper around an API. It is a layer between you and every AI tool that closes the gap between what you asked for and what you actually needed.
Is Kosmo an API wrapper?
No. An API wrapper sends your input to a model and returns the output. Kosmo does not replace your AI tool. It sits before it. Five agents work together to analyze, enrich, and restructure your raw intent into a compiled prompt before it ever reaches the model. It is a foundational multi-agent system, not a thin layer on top of someone else's API.
Does Kosmo guarantee perfect output every time?
No, and we do not claim that. What we see from our early users and our own testing is that compiled prompts almost always produce better output than uncompiled ones. We are still in public beta, actively tuning based on real user feedback. We have not published benchmarks yet because we want them to reflect our best work, not our first version.