A case study on designing Cairn, a budgeting and expense-tracking app. It turns transaction data into a visible trail of financial progress. Built for a 2026 world of AI-assisted, trust-first money management.
Try the clickable prototype ↓Cairn is a budgeting and expense-tracking app. It sorts spending into flexible categories and turns progress into a visible trail. It's built for people whose income doesn't arrive in tidy, predictable amounts. It's for people tired of red overspend banners that hide the real picture.
Underneath that trail is one simple idea. Every category and every goal is a stone in one ordered stack. They are not two separate features stitched together. Foundation stones are essentials, like rent and groceries. They sit at the base. Aspirational stones are savings and debt-payoff goals. They can only be added once the base is stable. Budgets and Milestones are just two views into that same stack.
Most budgeting apps still ask people to behave like accountants. Log every purchase. Force spending into rigid monthly categories. Absorb red, judgmental warnings when life doesn't cooperate.
Research into 2026 fintech behavior shows something different. Users don't quit budgeting apps because they stop caring about money. They quit because the apps make caring feel like homework. AI-driven "insights" often feel like surveillance, not support.
Design a budgeting experience with four goals. Remove manual tracking as the default. Explain every AI decision in plain language. Adapt to irregular income. Replace punitive spending alerts with a motivating, opt-in story of progress. Users still keep precise control over how much automation and personalization they want.
Three shifts define fintech UX this year. Explainable AI is table stakes now. It's no longer a way to stand out. Users will only link real accounts once trust is earned, before KYC. Financial wellness, more than transaction speed, is what keeps people coming back. Cairn was designed directly against these three shifts.
A five-stage double-diamond process, run end-to-end with weekly syncs alongside one PM and two engineers.
Research & competitive analysis
Personas & journey mapping
HMWs, IA & user flows
Wireframes to hi-fi UI
Usability testing & iteration
Five opportunities emerged for differentiating Cairn against the current budgeting app landscape.
Categorize spending by default. Let users correct it in one tap instead of building every category from scratch.
Every AI suggestion shows its confidence and its reasoning. Personalization is adjustable. It is never assumed.
Rolling budget periods and income-smoothing built for freelance and gig-economy earners. Their income never matched the fixed-salary assumption most budgeting apps still make.
A visual trail of milestones replaces red over-budget banners with empathetic, judgment-free language.
Most budgeting apps let you add a savings goal the moment you think of one. They don't check if your spending is under control. So the goal quietly gets abandoned the first time reality doesn't cooperate. Cairn works differently. It won't let a new goal sit above an unstable budget. Say you try to add Emergency Fund while Groceries is over budget. Cairn asks you to steady the base first. Or it flags the goal as at-risk from day one.
Understanding why budgeting apps fail to retain users, and what trust actually looks like in an AI-assisted financial product.
1:1 moderated interviews to surface the emotional and behavioral reasons behind budgeting-app abandonment.
A follow-up survey validated interview themes at scale.
n=112 per question. These are single-pick responses, so figures don't sum to 100. The remainder chose "other" or left the question blank.
A structured review of three category-leading budgeting products to map strengths, gaps, and openings for differentiation.
YBYNAB |
MMMonarch Money |
RMRocket Money |
|
|---|---|---|---|
| Description | Manual, zero-based budgeting method for intentional spenders | Net-worth & investment tracking with household collaboration | Bill negotiation and subscription cancellation, budgeting as add-on |
| Strengths | Deep methodology, strong community and financial education | Polished net-worth visuals, great for shared household finances | Effortless setup, genuinely useful subscription and bill tools |
| Weaknesses | Steep learning curve, entirely manual entry, no free tier | Premium-only, overwhelming for budgeting beginners | Upsell-heavy UI, budgeting feels secondary to monetization features |
| AI & Personalization | Minimal (manual control is the philosophy) | Moderate (automated insights, limited explainability) | Moderate (spend alerts, limited category flexibility) |
The gap in the market isn't another budgeting method. It's an app that combines the best of three: YNAB's intentionality, Monarch's visual clarity, and Rocket Money's low-effort setup. It drops their weaknesses too: manual burden, beginner overwhelm, and upsell fatigue. Cairn's chance is automatic categorization the user can trust, explained in plain language, wrapped in a progress story that feels earned, not gamified.
Synthesizing research into shared patterns, then into two personas that anchored every design decision.
"My worst month by the numbers was the month I turned down badly-paid work. The app doesn't know that."
"The app that gets close is the one that shows its work. The one that says 'trust us' is the one I delete first."
Scenario: Rohan sets up Cairn and reaches his first savings milestone.
| Phase | Awareness | Demo Preview | Account Link | First Week | Milestone Reached |
|---|---|---|---|---|---|
| Action | Sees Cairn recommended in a finance newsletter | Explores the app with sample data, no account required | Links his bank once he sees the security explanation | Reviews auto-categorized transactions in under a minute | Gets a quiet, tasteful notification: a stone added to his goal |
| Feeling | 🙂 Curious | 🙂 Reassured | 😌 Confident | 😊 Relieved | 🎉 Proud |
| Opportunity | Lead with outcomes, not features, in marketing | Let value be felt before any personal data is required | Explain exactly what is accessed and why, in plain language | Default to automatic categorization with one-tap correction | Celebrate without gamified noise. No confetti overload. |
Reframing research insights as "How Might We" questions, then structuring the product around them.
→ Persistent "why we flagged this" explainability tags, visible confidence scores, one-tap confirm or correct.
→ Replace red over-budget banners with a trail-and-cairn metaphor and empathetic copy ("15% over on dining" instead of "You overspent").
→ A demo-data preview mode lets users feel the product's value before any account is linked.
→ Rolling budget periods and income-smoothing suggestions instead of fixed calendar months.
→ A single "this week" home card with one clear, prioritized next action.
→ A settings panel that lets users dial automation from Manual to Assisted to Automatic.
→ One ordered stack of stones instead of two data models. Foundation-tier categories sit at the base. Aspirational-tier goals sit above. Goal creation is gated on foundation stability.
The "Trust" HMW earlier in this list isn't just a design principle. It needed real decision logic underneath it before any screen could show a number. "92% confidence" isn't a black box borrowed from a model card. It's a weighted match against the user's own transaction history. Every part of it is designed to be argued with.
| Confidence | Behavior |
|---|---|
| 90–100% | Auto-categorized with a one-tap Confirm. The user can raise or lower this line later. It's never fixed. |
| 60–89% | Suggested, but requires an explicit confirm before it's applied to the transaction. |
| 0–59% | Not guessed at all. Surfaced as "Needs your input" instead of a low-confidence label nobody trusts. |
Confidence for the old category drops for that merchant specifically. The corrected category becomes the new default, but starts fresh at 75% rather than snapping back to 92%. One correction isn't enough evidence to be that confident again.
Cairn stops auto-guessing that merchant entirely. Every future transaction from it is surfaced as a plain question. This continues until the user confirms the same category twice in a row. Three wrong guesses never happens, because the second one ends the guessing.
A correction only changes that exact merchant. Cairn never assumes that similar merchants, like other food-delivery apps, should change too. The user has to set a broader rule for that to happen. Past transactions are never silently recategorized. Only future ones are affected.
Budgets and Milestones aren't two separate data models wearing the same nav bar. Both tabs filter one ordered stack of stones by tier. A category (foundation) and a goal (aspirational) are the same underlying object, just filtered differently. A goal can't be promoted into the Milestones view until the foundation stones beneath it are stable. See the flow below, and the "foundation unstable" screen in the Design phase, for how that plays out.
Flow 1 · Reviewing an AI-flagged transaction
Flow 2 · Setting a savings milestone
From low-fidelity structure to a visual language built around trust, calm, and quiet celebration.
Low-fi stayed loose on purpose. The goal was to test the weekly-summary-first layout against the old dashboard-first mental model, before spending time on visuals. Mid-fi locked the layout once that held up. Then it added just enough color to test the stone-trail metaphor for legibility. Only hi-fi introduces the real palette and type. That way, early feedback stayed focused on structure, before aesthetics had a chance to distract from it.
Not every idea survived contact with users. The gamified home screen below is the clearest example. The reasoning behind cutting it shaped Cairn's tone more than almost any other decision.
Don't break the streak. Log a transaction today!
Gamified rewards weren't the top pick in the survey. Only 15% chose it, behind visible progress and personalized tips. But Revolut's savings-challenge streaks were a clear competitor benchmark. A daily streak plus a badge wall also tested well in isolation for "delight" during early sketch reviews.
The cairn milestone system kept the "visible progress" insight that made streaks appealing. But it tied that insight to a user-chosen goal instead of daily login pressure. Progress can pause. It never resets to zero.
| Style | Sample | Size · Weight |
|---|---|---|
| Display | Reach the next stone | 42–80px · 460 |
| Section head | Building the System | 30–42px · 460 |
| Card heading | A direction we explored | 18px · 520 |
| Body | You're 15% over your dining budget this month. | 16px · 400 |
| Data | ₹4,280.00 | 18px · 600 |
| Caption | Left to spend this week | 11px · 700 |
Deep pine and warm ochre replace the typical blue-and-white fintech palette. They feel grounded, not corporate. The rust accent is reserved only for genuine warnings, never routine spending alerts. That's how it keeps its meaning. Icons use rounded strokes to match Fraunces' soft serif terminals.
Warm serif plus an earthy palette is a familiar move by now. Plenty of fintech case studies reach for it to signal "not corporate." The part that's actually load-bearing here isn't the color choice. It's that every stone on the trail is a slightly different, asymmetric shape, not a uniform pill. No two are cut from the same border-radius. A real cairn works the same way: no two stones are ever the same.
High-fidelity screens from the final prototype, each tied to a research insight above.
Flagged because it matches 14 similar food-delivery purchases this month. Confidence: 92%.
Swiggy track record: 12/12 correct so far · overall accuracy 94% across 312 reviewed
You've placed your third stone. You've saved ₹15,000. Five more to go.
Cairn can see transaction history to categorize spending. It cannot initiate payments or transfers.
That's 18% more than your 3-month average of ₹7,140. Mostly weekday lunches.
Every element on the flagship screen traces back to a specific research finding.
Matches the 30-second check-in habit found in interviews. There is no full month of data to parse.
Replaces a bare completion percentage with a visual trail so progress feels tangible rather than abstract.
Highest-activity categories surface first, so the one thing worth reviewing this week is never buried.
The one category running hot uses the accent color. Color signals a decision point instead of decorating the screen.
All five destinations are reachable one-handed. See the reachability study below.
2026 fintech UX is judged as much on who it excludes as on what it enables. Three checks ran alongside every screen.
Primary actions, the weekly summary card, and the bottom navigation all sit in the bottom two-thirds of the screen. That's the zone a thumb reaches without a grip shift.
Before linking a bank account, Cairn shows read-only scope, encryption standard, and a revoke-access control on the same screen. None of it is buried in a settings menu three taps deep.
This isn't a screenshot. It's a working slice of the actual flow tested in the study below. Tap the flagged transaction, then confirm or edit its category.
Flagged because it matches 14 similar food-delivery purchases this month. Confidence: 92%.
Noted for next time. Future Swiggy orders will be categorized the same way automatically.
Cairn will remember this correction and apply it to similar Swiggy orders going forward.
Transactions the AI wasn't fully certain about surface first, with a confidence score attached.
Tap in to see exactly why a category was suggested. It's never a black box.
Either path teaches the model your preference for next time. No settings menu required.
Moderated testing to validate whether the explainability and control mechanisms actually worked in practice.
Round 1 tested 8 participants, 4 matching each persona, in moderated remote sessions via Maze. They worked through four core tasks. These were exploring the demo-preview flow, correcting an AI-categorized transaction, setting up an income-flexible savings goal, and adjusting the personalization level in Settings. After I shipped fixes for the findings below, Round 2 retested the same tasks with 11 new participants. This round ran unmoderated. It checked whether the fixes held up without me walking anyone through it.
Session recording · Task 2: correcting an AI-categorized transaction
This reaction showed a pattern: relief once the number was explained, paired with an instinct to distrust an unexplained label. It showed up in 5 of 8 sessions and directly shaped Finding 03 below.
6 of 8 participants didn't realize they could see why a transaction was flagged. The info icon only appeared on hover or long-press.
Made the confidence badge always visible instead of hidden behind an interaction.Several participants expected savings goals to live inside Budgets, since both showed progress bars that looked identical.
Gave Milestones a distinct cairn icon and stacked-stone visual language, separate from budget bars.Testers trusted "92% match" more than a plain "auto-categorized" label with no number attached.
Added a numeric confidence score to every AI-categorized transaction rather than a plain badge.5 of 8, then 10 of 11. These are different participants, so treat this as directional, not a controlled A/B.
The 8 and 11 participants above tested the happy path in depth. But every fintech flow eventually breaks in ways a moderated session won't surface on its own. A bank link times out. The AI has no confident guess. A category runs over. A goal gets added on top of an unstable base. I checked these five screens by inspection rather than moderated testing. Each one tests whether the same principles hold up under actual failure: explain, don't assume; calm over gamified; verifiable security; foundation before ambition.
Link an account or add your first transaction manually. The trail simply starts empty.
Your bank didn't confirm the connection in time. Nothing was read or stored. It's safe to try again.
₹1,860 over this month, still tracked in the same plain language. No rust in sight. That color stays reserved for account-security issues alone, even when this is the only category you have.
Emergency fund would sit above an unstable foundation. Steady Groceries first, or add the goal anyway and we'll flag it as at-risk.
Validated through iterative usability testing ahead of a planned engineering handoff.
Explainability only works if it's impossible to miss. Hiding trust-building information behind a hover state defeats its purpose. I also learned how much tone matters in financial messaging. The same underlying alert can read as supportive or judgmental depending entirely on word choice, and users notice the difference right away.
This is a self-initiated concept. Nothing here has shipped, so read the numbers above with that in mind. Participants were recruited informally through my own network, not a screened panel. The AI confidence-scoring in the prototype is simulated, not trained on real transaction data. Both testing rounds measured single-session comprehension, well short of the weeks of actual behavior change the app is designed to support.
Next, I want to test the rolling-budget model with a larger, more representative sample of gig-economy users. I also want to explore a lightweight voice interface for quick balance checks. Before any of this could ship, an engineer would need to validate the confidence-scoring model against real transaction data.