<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[Hamza's Blog]]></title><description><![CDATA[Hamza's Blog]]></description><link>https://topworkboot.hashnode.dev</link><generator>RSS for Node</generator><lastBuildDate>Mon, 05 Oct 2026 17:00:05 GMT</lastBuildDate><atom:link href="https://topworkboot.hashnode.dev/rss.xml" rel="self" type="application/rss+xml"/><language><![CDATA[en]]></language><ttl>60</ttl><item><title><![CDATA[How to Build an AI Application: From Use Case to Production]]></title><description><![CDATA[Building a demo with a language model takes an afternoon. Building an application that a customer pays for, relies on, and complains about when it breaks takes considerably longer, and the difficulty ]]></description><link>https://topworkboot.hashnode.dev/how-to-build-an-ai-application-from-use-case-to-production</link><guid isPermaLink="true">https://topworkboot.hashnode.dev/how-to-build-an-ai-application-from-use-case-to-production</guid><dc:creator><![CDATA[Hamza Majeed]]></dc:creator><pubDate>Tue, 08 Sep 2026 23:39:42 GMT</pubDate><content:encoded><![CDATA[<p>Building a demo with a language model takes an afternoon. Building an application that a customer pays for, relies on, and complains about when it breaks takes considerably longer, and the difficulty is not where most teams expect it.</p>
<p>The model is not the hard part. The model is a dependency you call over HTTP. The hard parts are choosing a use case where non-deterministic output is acceptable, building an evaluation loop that tells you whether quality is improving, and designing an interface that stays useful when the output is wrong — because sometimes it will be.</p>
<p>This is a practical walkthrough of that path, in the order the decisions actually arrive.</p>
<h2>Start with whether the use case tolerates being wrong</h2>
<p>The first question in any AI application is not which model or which framework. It is what happens when the output is bad.</p>
<p>Traditional software has a binary failure mode: it works or it throws an error, and you can test for both. Generative systems fail differently. They produce plausible, well-formatted, confident output that is wrong, and no exception is raised. Your architecture has to account for that, and if the use case cannot tolerate it, no amount of engineering will fix the mismatch.</p>
<p>A useful way to sort candidate use cases is by what the wrong answer costs and who catches it.</p>
<table>
<thead>
<tr>
<th>Use case shape</th>
<th>Cost of a wrong output</th>
<th>Who catches it</th>
<th>Viability</th>
</tr>
</thead>
<tbody><tr>
<td>Draft generation with human review</td>
<td>Low — the reviewer edits it</td>
<td>The user, immediately</td>
<td>Strong. Best first AI feature for most products</td>
</tr>
<tr>
<td>Classification into known categories</td>
<td>Low — one misfiled record</td>
<td>Batch review or downstream process</td>
<td>Strong, and measurable against labeled data</td>
</tr>
<tr>
<td>Summarization of source the user can see</td>
<td>Low — the source is right there</td>
<td>The user, if they check</td>
<td>Strong</td>
</tr>
<tr>
<td>Search and retrieval with citations</td>
<td>Medium — misleading, but auditable</td>
<td>The user, if citations are shown</td>
<td>Strong when citations are enforced</td>
</tr>
<tr>
<td>Autonomous action on external systems</td>
<td>High — the action already happened</td>
<td>Nobody, until later</td>
<td>Weak without confirmation steps</td>
</tr>
<tr>
<td>Advice in regulated domains</td>
<td>Very high</td>
<td>Often nobody</td>
<td>Avoid without domain review</td>
</tr>
</tbody></table>
<p>The pattern is clear: the viable early use cases are the ones where a human is already in the loop and the output is a draft rather than a decision. That is not a limitation to engineer around. It is the shape of nearly every AI feature that has actually worked in production software, and starting there is what lets you learn the rest. Sorting candidate use cases this way is the first step in our <a href="https://www.upsilonit.com/generative-ai-development-services"><strong>generative AI development services</strong></a> work, because rejecting a use case costs an hour and engineering around one costs a quarter.</p>
<h2>The architecture, concretely</h2>
<p>A production generative application is mostly ordinary software with a probabilistic component in the middle. The layers:</p>
<p><strong>Ingestion and retrieval.</strong> If your product needs to answer from your customer's data, this is the bulk of the engineering. Chunking strategy, embedding choice, metadata filtering, and the retrieval step itself determine output quality far more than model choice does. A mediocre model with excellent retrieval outperforms an excellent model with poor retrieval, consistently.</p>
<p><strong>Prompt and context assembly.</strong> Treat prompts as versioned artifacts in source control, not strings inline in a handler. You will change them constantly and you need to know which version produced which output when something regresses.</p>
<p><strong>Model invocation with a provider abstraction.</strong> Put a thin interface between your application and the provider from day one. Not a heavy framework — a small adapter. Providers change pricing, deprecate models, and have outages, and you want switching to be a config change rather than a refactor.</p>
<p><strong>Output validation.</strong> Never trust the output shape. Request structured output where the provider supports it, validate against a schema, and have a defined path for when validation fails: retry with a corrective prompt, fall back to a simpler prompt, or degrade to a non-AI path. This layer is what separates a demo from an application.</p>
<p><strong>Evaluation.</strong> A set of representative inputs with known-good outputs, run automatically on every prompt or model change, scoring results. Without this you are guessing, and you will make changes that feel like improvements and are not.</p>
<p><strong>Observability and cost tracking.</strong> Log every call with its inputs, outputs, latency, token counts, and cost. You will need this for debugging, for pricing, and for the conversation about why the bill grew.</p>
<p><strong>Caching, earlier than feels necessary.</strong> Two layers are worth having from the start. Exact-match caching on identical inputs is trivial to implement and eliminates a surprising share of traffic, because real usage is far more repetitive than test usage. Provider-side prompt caching, where available, cuts the cost of the large static portion of your context — system instructions, few-shot examples, retrieved boilerplate — which is often most of what you send. Both are far easier to add before the request path has grown complicated, and both change your unit economics enough to affect what pricing model is viable. Of the layers above, evaluation is the one Upsilon sees skipped most often and the one that makes everything else measurable.</p>
<h2>The trade-offs that decide the project</h2>
<p><strong>Latency is a product constraint, not an implementation detail.</strong> A generation that takes eight seconds is fine for a document draft and unusable inside a search box. Decide the acceptable latency before choosing an approach, because it eliminates options — long chains of model calls, large retrieval sets, and the largest models are all fast to build with and slow to run.</p>
<p><strong>Cost per request compounds in ways SaaS teams are not used to.</strong> Traditional software has near-zero marginal cost per action. AI features do not. A feature that costs a few cents per invocation is fine at a hundred users and a serious problem at a hundred thousand, particularly on flat-rate pricing. Model this before launch, not after.</p>
<p><strong>Fine-tuning is almost never the first answer.</strong> Teams reach for it early because it sounds like the serious option. In practice, better retrieval, better prompts, and better output validation solve most quality problems at a fraction of the cost and without locking you to a base model. Consider fine-tuning when you have exhausted those and have a substantial set of high-quality examples.</p>
<p><strong>Evaluation is the thing that gets skipped and the thing that determines success.</strong> It is unglamorous, it does not demo, and without it every prompt change is a coin flip. Build a small evaluation set before you build the feature, even if it is only thirty examples in a spreadsheet.</p>
<p><strong>Prompt engineering is not a substitute for a data pipeline.</strong> If the model does not have access to the right information, no prompt fixes that. Most "the model is not good enough" complaints are retrieval problems.</p>
<p>That evaluation loop is the part of the work that consumes the most time and generates the most value, because it is the only mechanism that tells you whether you are getting better or just getting different.</p>
<h2>A worked example</h2>
<p>We built an AI-powered proposal building platform that reached MVP in twenty-two weeks. That is long relative to a typical MVP, and every additional week went into the same place: the loop between generated output and acceptable output.</p>
<p>The engineering around the model was not the constraint. Ingestion, retrieval, the editor, and the export were conventional software and moved at conventional speed. What could not be compressed was finding out whether the generated proposals were good enough that a professional would send them to a client. That question has no answer in advance. It only has an answer after running real inputs, reading the output, adjusting retrieval and prompting, and running it again.</p>
<p>Two decisions made that loop tractable.</p>
<p>The first was scoring output on a small fixed set of representative inputs rather than on whatever came in that week. Fixed inputs mean changes are comparable. Without that, every improvement is anecdotal and the team argues from impressions.</p>
<p>The second was designing the interface around editing rather than acceptance. The output landed in an editor as a draft, with the retrieved sources visible alongside it. That framing meant a mediocre output was still useful — the user edited it, which was faster than starting from nothing — and a wrong output was obvious rather than hidden. It also generated exactly the signal the team needed, because what users edited told us where quality was weakest.</p>
<p>Everything else was cut to protect that loop. One user role. No team management. Billing by invoice, outside the product. Templates maintained by hand. All of it added later, once the core question had an answer.</p>
<h2>A decision framework before you write any code</h2>
<p><strong>Write the failure sentence.</strong> "When this produces a wrong output, the user will [notice how] and [do what]." If you cannot complete it, the use case is not ready.</p>
<p><strong>Build the evaluation set first.</strong> Twenty to fifty representative inputs with what a good output looks like. This takes a day and changes every subsequent decision.</p>
<p><strong>Set the latency and cost budget explicitly.</strong> Milliseconds and cents per request, written down. They eliminate architectural options immediately, which is the point.</p>
<p><strong>Assume the model will be swapped.</strong> Providers deprecate, prices change, and better options appear. A thin abstraction now costs an hour; retrofitting it costs a sprint.</p>
<p><strong>Design the interface for the wrong answer.</strong> Show sources. Make output editable. Add confirmation before consequential actions. The interface is where trust is won or lost, and it is not a model problem.</p>
<p><strong>Instrument everything from the first call.</strong> Inputs, outputs, latency, tokens, cost, and user action taken afterwards. That last one is your real quality signal.</p>
<h2>Common questions</h2>
<p><strong>Which model should I start with?</strong> Start with a capable general-purpose model from a major provider, get the pipeline working, then evaluate alternatives against your evaluation set. Choosing a model before you can measure quality is choosing without information.</p>
<p><strong>Do I need a vector database?</strong> Only if you have enough documents that filtering and ranking matter. Below a few thousand chunks, simpler approaches are often fine and considerably easier to debug. Add the infrastructure when the simple version breaks.</p>
<p><strong>How do I stop it from making things up?</strong> Constrain it to retrieved context, require citations, validate output structure, and show the user the sources. You reduce the rate and make the remainder visible; you do not eliminate it, and a product design that assumes you can is fragile.</p>
<p><strong>How much does an AI MVP cost compared to a normal one?</strong> Similar for the surrounding software, plus the evaluation loop, plus ongoing inference costs that traditional software does not have. The recurring cost is the part most budgets miss.</p>
<p><strong>Should I use an agent framework?</strong> For a first production feature, usually not. Frameworks add abstraction over a problem you do not yet understand, and debugging a multi-step agent chain is significantly harder than debugging a single call with good logging. Start simple and add structure when the simple version demonstrably fails.</p>
<p><strong>How large does the evaluation set need to be?</strong> Twenty to fifty inputs is enough to detect meaningful regressions, which is what it is for. It does not need statistical rigor; it needs to be fixed, representative, and run on every change. Grow it whenever a production failure surfaces a case it did not cover.</p>
<p><strong>Should the AI feature be optional for users?</strong> Early on, yes. A product that works without the AI feature and is better with it degrades gracefully when quality is poor or the provider is down. A product where the AI is the only path has no fallback, and outages are not hypothetical.</p>
<p><strong>How do I price a feature with real marginal cost?</strong> Measure cost per request from the first day of production, then decide. Usage-based or credit-based pricing aligns cost with revenue; flat pricing requires a cap or a routing strategy that keeps the average down. Choosing before you have the data is guessing at your own margin.</p>
<h2>The short version</h2>
<p>Pick a use case where a human reviews the output. Build the evaluation set before the feature. Treat retrieval as the main engineering problem, put a thin abstraction in front of the provider, validate every output against a schema, and design the interface so a wrong answer is visible and editable rather than hidden. The model is the easy part.</p>
<p>If you want a second set of eyes on an AI use case before committing engineering to it, Upsilon has been building software for early-stage teams since 2012.</p>
]]></content:encoded></item></channel></rss>