by ConKarma team · 2026-05-15
Why ConKarma's conflict tools aren't AI-mediated
We built sibling, couple, and cross-generational repair flows out of curated content and structured forms — never an AI in the middle. Three reasons: privacy, calibration, agency. Here's the long version of each.
The kid is crying. The other kid is screaming. You're in the kitchen, the spaghetti's overcooked, and the eight-year-old has just told the six-year-old she's adopted (she's not). You'd like a tool that helps the next ten minutes go better. Most apps will offer to do that with an AI: paste the situation in, get a script back, hand it to the kids, watch them parrot the script with the same tone they use to read aloud.
We chose not to build that. ConKarma's sibling, couple, and cross-generational repair flows are curated content and structured forms — never an LLM in the middle. This is a deliberate, durable position. Three reasons.
The market default
Open any "AI for families" pitch from 2024 onward and you'll find the same pattern. Conflict → describe situation → AI generates message → user sends. It's a tidy product surface, easy to demo, and it sells well to investors. It also has three problems that compound the longer you use it.
The first is what you give up to make it work. The model needs context. Useful context means actual content: what the kids said, what they did, what the room felt like, what the prior arguments have been about. To get a useful response, you have to send that material somewhere — your phone, then a server, then a model. Even with the best privacy posture you can offer (no training, no retention, no logs), you've routed the most intimate operational data of your household through somebody else's pipeline. Once you've done that, you can't un-do it. The next argument carries the same exposure.
The second problem is calibration. LLM outputs are confidently wrong, in two specific ways that matter for repair. They overshoot — escalating "I'm frustrated" into language that lands like accusation. And they smooth — flattening real frustration into pleasant-sounding script that the receiving partner immediately notices is not how this person talks. Neither failure mode is debuggable from the user's side. The model is non-deterministic; the same input on a different day produces a different output. There's no way to learn the tool's mistakes and adjust your trust in it. You're guessing each time whether what came back is the right register.
The third is agency. When you outsource the script, you outsource the practice. A repair tool that hands you words is doing the most important work for you — the work of finding the words. The argument with your sibling that you actually remember decades later is the one where you had to figure out what to say. We don't want to shortcut that for the kids. We want the tool to scaffold them through the structure (acknowledge, name, repair, agree) without picking their words for them.
Our position
We didn't ship the AI version of any of this. What we shipped instead, for the three repair surfaces in ConKarma:
Sibling repair flow. Each kid sees a curated set of "what happened" cards — short, age-appropriate phrases written by editors, not generated. The first kid taps the ones that match what they remember. The other kid does the same, on their device, without seeing the first kid's selections. The BE matches the pair into a shared prompt — a single screen that surfaces what both kids agreed on, what only one of them said, and a curated "what to say next" card. Both confirm before the flow closes; either can step back to the picker. The curated set covers the patterns siblings actually argue about — possession, fairness, attention, accidental harm, betrayal of a confidence — at three reading-age tiers per reading-age-aware copy, so a six-year-old and a fourteen-year-old see the same flow with vocabulary they can actually parse.
Couple repair flow. Longer cadence than the sibling version because adults can take longer to cool down. The partner who starts the flow drafts a framed message from a fixed prompt library (about thirty templates we've editorial-reviewed, organised by the underlying ask — I need to hear, I want to acknowledge, I want to ask for, I'd like to try again). Optional 24-hour cool-down before the second adult sees it: if you start the flow at midnight after a fight, you can choose to queue it for noon the next day. The receiving partner sees the message in a frame that includes the original prompt template — so they're reading not just "you said X" but "you reached for the I want to acknowledge template, and inside that, you wrote X." That metadata matters. It signals which conversation the sender thought they were having. Both sign off before a "moment closed" badge appears.
Cross-generational repair flow. Same structured-form discipline, but with role-aware prompts. A parent reaching to a teenager picks from a different template set than a teenager reaching to a parent; an adult-child reaching to an aging parent in the elder cell picks from a third. The prompts know who's reaching to whom, and adjust the entry surface accordingly. The Family Shield routing per /help/family-shield sits underneath all three repair flows — when the kid's input matches one of the trained crisis keywords (self-harm, abuse, danger), the prompt set immediately re-orients to surface a local resource hand-off instead of continuing the repair flow.
All three flows share five properties:
- Every text the user sees comes from us, not a model. The phrase the kid taps is a string we wrote and edited. The template a partner picks is one of thirty we authored. The "what to say next" card is one of a hundred we reviewed.
- The user does the writing where writing is required. The couple flow includes free-text fields. We don't auto-fill them, don't suggest completions, don't post-process. The words you write are the words your partner sees.
- The flow is deterministic. Same selections produce the same screen. There's no model variation to debug.
- The data stays in the cell. Sibling-repair selections, couple-template choices, cross-gen prompt picks — none of it leaves the cell's encryption boundary. No analytics record beyond "a repair flow occurred" (and even that doesn't carry the template, just the surface).
- The Family Shield hand-off has no AI in it. It's a keyword match against a small, hand-curated trigger set. False positives go to a human-reviewed queue. False negatives are a known limitation we don't try to AI our way out of.
When AI is the right tool
We're not anti-AI. ConKarma uses Claude in surfaces where the failure mode of getting it wrong is mild — generating the monthly family-narrative recap from anonymised aggregates (in-app only, opt-in, /help/ai-privacy), producing the personal coaching note that lives only on your own device, drift detection on engagement patterns (no content), affiliate-curation recommendations. In each of those, if the model overshoots or smooths, the user can see what happened, ignore it, and move on. The cost of the bad output is one paragraph of bland copy, not a sibling escalation.
Conflict is the inverse case. The cost of a bad output is high, the calibration window is narrow, the user has no way to debug, and the data that would be needed to get it right is the most intimate data the household produces. That's the wrong shape for an LLM call. It's the right shape for editorial work, structured forms, and Family Shield human routing.
Why this is durable, not transitional
It would be reasonable to read this as a current position — a "we're not ready yet" stance that will flip when models get better, more private, or cheaper. We don't think so, for two reasons.
The first is that the privacy posture is structural, not technical. Even with on-device models, even with zero-egress guarantees, even with verified-no-training, the act of routing intimate content through a model is a different kind of operation than not routing it. The model is a third party in your kitchen. Some users will be comfortable with that; some won't; both are reasonable. Building the product so it doesn't require routing through a model means we don't have to defend a position one camp will always disagree with.
The second is that the calibration problem is not a model-quality problem. Even a perfectly-calibrated model would produce a generic-feeling output, because the inputs it received are common across millions of households. The output that lands is the one that sounds like you — your sibling, your spouse, your kid. The path to that output is for you to write it. A tool that scaffolds you to the writing is doing more useful work than a tool that drafts the writing for you.
What this looks like in the product
If you're a user of ConKarma, the visible difference is small. You'll see a Hard moments section on /features with curated card sets, you'll see a "draft a message" flow on couple-repair that asks you to pick a template before you write, you'll see prompts that are clearly written by people not generated by a model. The product feels editorial and structured rather than chatty and helpful.
If you're a press reader trying to slot us into a category: we're the family + couple software company that turned down the AI-mediation route. We make different tradeoffs from the LLM-default cohort. The trades show up in the /security posture, in the /security/zone-isolation guarantees, and in the way the in-app surfaces feel.
If you're a clinician or researcher: the repair flows aren't a substitute for the work you do. They're a scaffold for the kinds of small moments that happen between sessions — the eight-year-old / six-year-old moments, the post-disagreement check-ins that don't need a fifty-minute slot, the cross-generational asks that wouldn't make it onto your agenda anyway. The deeper work still belongs to you.
What we got wrong, what we don't know yet
A few honest notes:
The curated-content corpus is the bottleneck. Editorial-reviewing thirty couple templates × three reading-age tiers × fourteen locales is real work. We will be slow to grow it. That's a feature in the privacy and calibration sense, and a bug in the breadth-of-coverage sense — there will be moments the templates don't quite fit. The escape valve is the free-text field: you can always write what you mean.
The Family Shield routing is a keyword match against a small set we trained editorially. It will miss things. It will catch things that aren't crises. We don't lean on it for safety guarantees we can't keep. The clinician hand-off is one tap; the resource list per /help/family-shield is curated and locale-aware.
We haven't fully measured whether the curated-prompt format works better than an AI-generated alternative. We have anecdotal feedback from the closed beta and we're tracking aggregate engagement (does the flow get completed, do users come back). We're not running A/B tests against an LLM arm because we don't want to ship one even temporarily.
Closing
Conflict tools are not a place to be clever. They're a place to be deliberate. We chose curated content and structured forms because they preserve the household's privacy, produce calibrated outputs, and leave the agency of finding-the-words with the people doing the finding. Three reasons, one position, durable not transitional.
If you've been waiting for a family app that doesn't route your hardest moments through somebody else's model — that's the one we shipped.