Skip to content
Bernhard Götzendorfer
Behind the Scenes

HACK_002 in Vienna: Pflegegeld-Prüfer wins Track A1

HACK_002 recap from Vienna: the Pflegegeld-Prüfer, a care allowance checker, wins Track A1 but misses my advance target of 12 of 14 court cases. It gets 10.

Bernhard Götzendorfer at his table at HACK_002 on Saturday, laptop open, next to crisps, a can and a water bottle.

TL;DR

At HACK_002 at the HOIV, the Home of Innovation in Vienna, I built the Pflegegeld-Prüfer. It's a care allowance checker, and it won Track A1. You describe a normal care day, and the checker works out whether the level in the official decision fits. The target I set in advance was 12 of 14 court cases. It got 10. Both numbers are on the website. It is live at www.pflegegeld-pruefer.at, free and with no ads.

Submitted while I was asleep

Sunday, 04:54. The last pickup for the submission video is in the can. Two minutes later I approve the submission in writing. Then I go to bed.

At 05:58 an agent sends the submission: ZIP, video, form. I was already asleep. The deadline would have been 08:30.

At noon I demoed the checker live. Track A1, "Applied AI for Consumers", went to the Pflegegeld-Prüfer. 65 people in 25 teams took part. To be honest, I'm really glad it was enough without me moving the target afterwards.

And the submission video had sound. At my first hackathon in March, that is exactly what went wrong.

Rosa, 81, and three jumpers

In Austria, the care allowance, Pflegegeld, comes in seven levels. Which level someone gets depends mostly on the hours of care needed per month. If you are one level too low, you lose money every month. From level 3 to level 4, that is currently 295.90 euros.

Does that happen often? Beforehand, my agents went through 29 published court cases that include the original decision. In 13 of them, the court set a higher level than the first decision. In none of them a lower one. That doesn't tell you how often the authority gets it right. Only disputed cases go to court, and only some of those get published.

The checker turns the form around. Instead of ticking boxes, you describe a normal day, the way you actually talk. Rosa is 81 and invented, like all the example families. Her daughter says, in Austrian dialect: "In the morning I have to stand next to her, or she puts on three jumpers." A language model maps each sentence to a rule from the law or the Austrian classification regulation. It has to quote the sentence word for word. The code checks every quote. A line without a verbatim quote does not count.

All the math happens in code: hours, level, euros per month and the deadline for an appeal. Not a single number comes from the model. The claim on the start page puts it in two sentences: "Tell, don't tick. AI listens, code does the math."

Two screenshots for Rosa: level 4 instead of 3, about 295.90 euros more a month, and her sentences with the rule

I only measured dialect on the side. Seven invented cases were rewritten in dialect, and none of them changed level. That was not preregistered, so it is just a side note here.

The result is an estimate, not a decision. It prepares a counselling session or the next assessment, and it is not legal advice. Here is what it looks like in 54 seconds:

The Pflegegeld-Prüfer in 54 seconds, with Rosa as the invented example case. On-screen text in English, no voice. The background music is AI-generated.

22 minutes early

In my pre-event post on 23 September I wrote that there would be no repository before the opening. I need to correct myself on that.

The repository existed from Saturday 08:25, with documents only and not a single line of code. Code writing started at 11:45, shortly before the opening. The first code commits are stamped 12:38. The official start was 13:00. The submitted README said so. In the end the ZIP went out without the git history, so the times were in the README and in the log. I rewrote nothing.

Whether prepared code was allowed was one of my three questions for Sunday. I never got it answered. So the deviation is stated openly, instead of me explaining it when someone asks.

The idea was also fixed earlier than announced. I did not choose at the opening. I chose on Friday at 15:26: care allowance. By 17:34 it was sharper: not estimating the level, but recalculating the decision. A decision has an hour count, a level and a deadline. That can be recalculated.

One point of contact, many agents

I built solo, with agents. Three sessions each took one area: calculation core and measurement, mapping with AI, and the interface. A coordinator session kept interfaces, test cases and the README together. On top sat an observer session, and that was pretty much the only one I talked to. I tested, gave feedback and tested the next version again. Direction, naming and what the checker is allowed to claim stayed with me.

Merges happened only on the hour and the half hour. Even so, main went red twice. Once from 15:34 to 16:00, once from 17:02 to 17:11. Both times the individual changes were green on their own, just not together. From 17:05 on, one session checked main plus all candidates together before each window.

In the pre-event post I announced points where I would cut scope instead of building on. There were two, 16:00 and 23:00. The project passed the first at 16:02. The name came at 17:00. The technical working title became Pflegegeld-Prüfer, because the target group should understand what it is before they click. I signed off the second point at 22:37, with the complete flow running in the browser.

By submission there were 885 commits in the history, 330 of them merges, plus 142 issues and 205 merged merge requests. The first code commit and the last commit are about 15.5 hours apart. That is shared work with agents, documentation included.

17:47: Rosa lands one level too low

At 17:47 a helper session reports: through the interface, Rosa lands at 153 hours. That is level 3. Her invented case is set up for 163 hours and level 4. With the full transcript all five example cases hit. Through the interface, Rosa was ten hours short.

My decision at the time: the follow-up questions aim at the hours missing for the next level. If needed, there is a second round with at most six questions. And the display never goes below the decision. A result below the decision would send the wrong message to a family.

After the change, Rosa reached 163 hours after five questions in two rounds. Level 4.

The number, as promised

In the pre-event post I said I would show the evaluation after the event. Here it is.

The target went into the repository on Saturday at 12:38, with the first commits. That was four and a half hours before the first run of the app. At least 12 of 14 published court cases should hit exactly the right level. Right mostly means the level from the ruling. For children, right means giving no level at all, just a note. There was an earlier preregistration on Friday evening, but it only covered the model comparison.

The first app run at 17:13 got 9 of 14. At 19:01 I froze the rules and the prompt. From then on nothing was tuned on these cases. Then the final run, three runs per case: 10 of 14. The 95 % interval is 45 to 88 %. Target missed.

Two of the four misses are too low, because one item stayed open. Once the checker failed to recognise a child as outside its scope. And then one case, blind and with an amputated leg, level 4 according to the court. It ends at level 1 in all three runs, because the checker misses the minimum level for blindness. On these 14 cases it was never too high.

If you count whether it asks the decisive follow-up question, 13 of 15 cases are solved. The target for that was 13, and it was met.

Then the second test set with 15 more cases, measured once after 22:30: 5 of 15, interval 15 to 58 %. Clearly worse. Twice it was too high there. And on these cases the original decision was right more often than my checker, 8 of 14. It wasn't a blind test. The cases had already run in the model comparison, and some of the rules come from them.

The baseline is the original decision. It matched the level from the ruling in 16 of 29 cases, and all 13 misses were too low. That is only partly comparable.

On Monday I also measured the model that runs on the website. Gemini 3.8 Flash got the same 10 of 14, with the same four misses. Had I written the target down only after the run, it would probably say 10 now. That is exactly why you write it down first. Every case, misses included, is on the methodology page.

After midnight: freeze, and then not quite

After the second point at 22:37, the question from the pre-event post was on the table: cut or keep building? I chose to keep building. Among other things, a card was added where you confirm or deselect uncertain items.

At 00:48 the freeze was in place, with a last merge window at 01:30. At 02:12 I lifted it again, as far as it helps the app. The last commit is stamped 04:12. That went against my own plan, and that is exactly how the log records it.

In between, the video came together. "For the video I also took a hundred percent agentic route," as I put it on Sunday. Agents built the slides, screenshots and the cut. I recorded screen and voice at 03:25 and did pickups at 04:49.

Two winners at one table

I registered solo and ended up at a table of four, with Helena, Hayat and Julian. We brainstormed at the same table, talked and spent the evening together. Each side built its own project. Theirs is called miRacle and shared the win in Track B3. Two winners at one table, I like that a lot.

There was room for an evening like that because I didn't steer every session myself. I tested and cleared bottlenecks. The observer session handed out the work.

Three questions, three answers

I took three questions with me on 23 September.

Did I miss anything in the night that I would have been allowed to prepare? Yes, two things. The shared test before merging only came after main went red the second time. And I had no contact with a real person from the target group. I did not manage the user test that Track A1 actually asks for. The evidence rests on court cases and invented families, and that is exactly what the submission says.

Can you get a measured result with a baseline into 19.5 hours? Yes. A small real number with an interval, and a baseline from other people's decisions. The claim is no bigger than that, and nothing about it is clinically validated.

What does the opening say about prepared code? I never found out. My answer is the open timeline in the submitted README.

Sunday afternoon: the Pflegegeld-Prüfer goes live

That same afternoon the checker went live on its own domain. It came with an imprint, a privacy policy and a consent step before your own text goes to the AI service. The texts are now written for real visitors.

The next step is a field test with a care counselling professional and with family carers. Until then the checker stays what it is: preparation for a conversation.

Hack-Nation 7 is on 3 and 4 October, again at the HOIV, and I'm registered for the Vienna hub. Coming with me: the shared test before every merge, a single session as point of contact, and the preregistration. The Pflegegeld-Prüfer code stays at home. Only what I learned comes along. And if I start early, this time I will say so beforehand.

If you know someone in Austria who has a care allowance decision in front of them right now, pass the link on. How I work with agents day to day is on the start page. And if you have feedback on the checker, write to me.