Skip to content
Bernhard Götzendorfer
AI Deep Dives

From 250+ Prototypes to Product: What I Learned

What sits between an AI prototype and daily use: data, failure handling, ownership, and deciding when an experiment has answered its question.

TL;DR

I have worked with AI agents every day since late 2024. Across more than 250 of my own prototypes and experiments, one distinction has stayed useful: an experiment answers a question. A product also needs a way to operate day to day. This article covers the work in between and how I decide when an experiment can stop.

What question should the prototype answer?

For a prototype, I start with a limited question. Can a model read the required fields from a receipt? Can it find the relevant information in my documents? Can it prepare a recurring work step?

One useful result says little about how the workflow handles other inputs. Difficult examples belong in the early tests too: a badly photographed receipt, an incomplete record, or a request for information that is missing.

If you have not chosen a use case yet, the getting-started guide for businesses describes how to select a single process.

The work around the model

Receipt recognition was an important part of building BuchhaltGenie. Alongside the model, I had to work on the inputs and what happened to the output: different document formats, preprocessing, and structured handling of the results.

A complete workflow raises further questions:

  • Inputs: Which formats arrive? What happens when something is missing or unreadable?
  • Review: How do I recognise a wrong answer? Which results need approval?
  • Integration: Where does the result go, and who may access it?
  • Failures: How does a failed step become visible? Can I retry it without creating duplicates?
  • Operations: Who responds when quality drops or an interface changes?

A different model may help. The surrounding work remains part of the project.

Small changes I can check

In my development sessions, I work on bounded changes. I check the result before adding the next step. For AI output, that requires concrete examples with expected results. An answer that merely sounds plausible is hard to compare.

I also record why I chose or discarded an approach. When I return to a project, I need that decision and the open questions. The next session should be able to continue from there.

When the experiment is enough for now

An answered question does not always need to become a product. I look at four things when making that decision.

Does it work with the actual inputs? If the approach only holds with selected examples, I need more experiments before expanding the workflow.

Does the effort make sense? Model costs sit alongside integration, review, and maintenance. I work with the expected volume and the effort the task currently takes.

Is failure handling clear? An incomplete draft can be corrected. A wrong action in a connected system may create more work. Approval needs to fit the particular step.

Who looks after it once it is running? Even in my own projects, this is a question of capacity. Another service needs time for updates and failures.

If these pieces do not fit yet, the prototype can wait. I record the question it answered and the limits of the approach.

If it should become a product

I would plan the transition in four steps. Their duration depends on the particular project.

  1. Check real examples. Collect expected results, failure cases, and running costs. Define in advance what quality is useful for the task.
  2. Prepare for operation. Work through access, data storage, failure handling, monitoring, and the relevant legal requirements.
  3. Start with a limited scope. Use the workflow at a manageable volume, collect feedback, and continue checking its results.
  4. Record ownership. Document who handles updates and failures, and how to disable or roll back the feature.

The prototype helps me decide whether this further work is worthwhile. After that, I plan the scope for daily use.

If you are making that decision about your own prototype, you can describe the workflow through the contact section.