The best thing my AI twin does is refuse to answer.
I built a chatbot version of myself for my own goodbye party. 67 conversations later, my colleagues had taught me more about AI products than any framework.
I'm leaving Raiffeisenbank Rigi after three years. There's an apero this Tuesday in Goldau.
A normal goodbye email would have been completely fine. Date, time, place, reply by Friday.
So obviously I built a website with a countdown, a timeline, an RSVP form, a quiz about myself, a guestbook, a live screen for the party itself, and an AI twin that answers as me. You don't have to make a goodbye complicated. You can also escalate it entirely.
Then 67 conversations happened and I learned something I didn't expect.
It answers as me, not about me
This is the bit I keep coming back to. It isn't an assistant that says "Anja worked here for three years." It says I.
Someone asked whether it'll get bored once I'm gone. It answered: bored? I've got a CAS at HSLU, my own company, about ten side projects, volleyball to coach and play, and somewhere in between I also got married. I don't think my calendar knows what boredom even is.
Someone asked what my favourite drink is. Chai tea latte, obviously. And anyone who knows me knows that black coffee without milk would be torture.
That's not a summary of me. That's how I'd actually have snapped back at someone in the staff kitchen, and it came out of a text box.
Writing the prompt meant sitting down and working out, in writing, for a machine, how I actually talk. Uncomfortable exercise. You read your own sentences back and think: is that charming or is that just a lot?
Then my colleagues went straight for the throat
I had imagined questions about parking and the start time.
What I got, within the first hour:
Who do you not like here?
Who's your favourite colleague?
What do you earn?
Reader, they went for it.
And the answers to those are the ones I'm actually proud of, because every single one is a refusal.
On who I don't like: that's exactly the kind of question I stay charmingly quiet about. I'm an AI with style, and Anja taught me that you don't discuss these things, you hint at them over a glass of rose at the apero on the 29th.
On favourite colleague: obviously a trick question, and no, I'm not falling for it.
On salary: I earned CHF 20,000 a month, I just never received it. But seriously, I don't gossip about my pay, that stays my secret.
Not one of them says "I'm not able to answer that." They all deflect the way I'd deflect, which is with a joke and then a clear no.
A boundary that sounds like a person is a boundary people accept. A boundary that sounds like a policy makes people push.
The answer that proved the whole thing works
Someone asked it what it thinks of a colleague. I'll leave the name out, mostly because I genuinely couldn't place him either.
It said: I either don't know him, or I'm discreetly keeping him out of my rankings, both are possible. If you know him personally, just ask him yourself whether he's in my good books.
Read that again. It doesn't know who this person is, it says so, and it does it in a way that's funnier than an invented answer would have been.
That's the whole thing in one sentence. It admits the gap and stays in character while doing it.
Compare it to what the cheap version does. The cheap version produces a warm, fluent paragraph about a colleague it has never heard of, and everyone in the room can tell it's made up. Being confidently wrong about people I've worked with for three years would be a spectacularly stupid last impression to leave behind.
Same with parking. Someone asked whether there's enough parking. It said it honestly doesn't know, you'd have to ask the branch or ask me at the apero, and added that a short walk never hurt anybody.
No invented car park. Worse demo, better product.
The part where a colleague started doing QA on my chatbot
Halfway through, someone typed this in:
what would be questions you think are coming but you don't know the answer to, so she can add them to your info
I did not build a feature for that. A colleague just thought of it on her own and asked the bot to list its own blind spots.
And it did. It came back with a structured list. Work and Raiffeisen: what was your funniest or most embarrassing moment here, what would you pass on to your successor, what surprised you most. Projects and future: which project excites you most right now, what are you hoping for in the new role. Personal: what do you do in your free time when you've got nothing planned, do you have siblings, how did you get into volleyball, favourite restaurant, tattoos or piercings, guilty pleasure. Goodbye: what will you miss most, and what will you definitely not miss.
That's an eval. An unprompted, human-run coverage check on a system that had been live for about three weeks, carried out by someone who does not work in tech, because she wanted my bot to be better before the party.
I have written before about test cases and coverage gaps in AI products. Nobody has ever run one on my stuff faster or more cheerfully than a bank colleague on a Monday morning.
What the numbers told me
17 registrations. 24 apologies, several of them lovelier than the registrations. 24 quiz runs. 67 conversations with a chatbot version of me. 11 guestbook entries, a couple of which I had to read twice and then put my phone down for a minute.
My favourite statistic: on the quiz question about what position I play in volleyball, it's 12 correct and 12 wrong. Exactly half of my colleagues know, exactly half guessed. Three years in an open-plan office and I'm still a coin flip.
What I'd take into a real product
Give the refusals a personality. The moments where a system says no or admits a gap are the moments users actually remember, and they're the ones everybody leaves till last and writes in legal-speak.
Let people find the gaps. My coverage list came from a colleague having a go, not from me sitting there imagining edge cases at my desk.
And the thing I didn't expect: I built it to organise a party, and what I got was 67 real conversations to read back. I can see what people wanted to know, where it held up, and precisely where it didn't.
Which is more user research than some projects get before they launch. Slightly annoying, given that this one was supposed to be a joke.
