AI Financial Advice Is Better Than You Think. That's the Problem.
For years, the standard answer to “should I ask an AI about my money?” has been a nervous no. Hallucinations plus tax law equals disaster, the thinking went. Recent work out of MIT Sloan lands somewhere else entirely. The answers held up fine. The problem lives somewhere the accuracy debate never looked.
A quick note on where the conversation isn’t
Worth saying upfront: this topic is oddly quiet online. Sweep the last month of Reddit, Hacker News, and X, and you’ll find almost nothing on the quality of AI financial advice. Meanwhile the same communities generate hundreds of daily threads about AI coding tools, benchmark scores, and which model writes better React.
That silence is itself interesting. People will publicly hand their codebase to a model and argue about it for 400 comments. Almost nobody publicly admits to asking one whether to max out their 401(k) or pay down the mortgage. They’re either doing it quietly or avoiding it quietly. So this piece isn’t a roundup of community sentiment. It’s about the structural problem the research points at.
The answers were fine
The conventional worry made sense. LLMs generate confident-sounding text, and personal finance is full of hard numbers — contribution limits, marginal rates, account eligibility rules — where a plausible-sounding wrong answer costs real money.
Then you actually run the standard questions. Traditional versus Roth. Which debt to attack first. How many months of expenses to hold in cash. Whether to buy individual stocks or index funds. Current frontier models answer these correctly, and often better than the average financial content on the open web.
The reason is unglamorous. Personal finance fundamentals are the most repeated advice on the internet. Pay off high-interest debt first. Don’t time the market. Watch the expense ratio. Capture the employer match. That corpus is enormous, consistent, and largely correct. A model trained on it is going to nail the basics.
There’s even an argument the AI has a structural advantage over a human. A commission-based advisor gets paid differently depending on what they recommend. In the US that conflict is old enough to have its own regulatory history — the whole fiduciary-standard fight exists because “suitable” and “best for the client” aren’t the same thing. The model has plenty of failure modes, but selling you a high-load annuity isn’t one of them.
The bottleneck is the question
Here’s where the research gets uncomfortable. The variable that mattered wasn’t answer accuracy. It was question quality.
Ask a model: “How should I save for retirement?” You get boilerplate. Start early, diversify, take advantage of tax-advantaged accounts, consider your risk tolerance. Every word is true. Almost none of it is useful.
Now ask this instead: “I’m 39, earning $145K, with $310K left on a mortgage at 4.2%. I have $180K in a 401(k) with a 4% employer match, currently maxing neither that nor an IRA, and about $1,400 a month in free cash flow. On an after-tax basis, does that money do more work as extra mortgage principal or as retirement contributions?”
Now the model does actual math. It weighs the guaranteed 4.2% after-tax return against expected market returns. It flags that the employer match is free money and should come first. It asks whether you’re itemizing. It walks through the tradeoff.
Same model. Radically different value. And the gap between those two prompts has nothing to do with AI capability — it’s a gap in what the asker already knew. To write the second question, you need to know a 401(k) match exists, that mortgage prepayment is a risk-free return, that comparisons belong on an after-tax basis, and that your own numbers are the inputs that matter.
So the gap widens instead of closing
The optimistic case for AI financial advice was access. Comprehensive planning has historically been for people with enough assets to make an advisor’s fee worthwhile — the households that need help least. A free, always-available advisor sounded like a leveling tool.
The research points close to the opposite. AI rewards people who ask good questions, and asking good questions requires the exact domain knowledge that separates the financially literate from everyone else. A CFA-adjacent professional uses a model as a calculator and a sanity-check, multiplying expertise they already had. Someone starting from zero gets generic advice and, critically, has no way to evaluate whether the generic advice fits their situation.
The failure mode gets worse than that. Models inherit your premises and answer them faithfully. Ask “is whole life insurance a good fit for me?” and you’ll get a diligent, balanced breakdown of whole life insurance — even if you’re a 28-year-old with no dependents who has no business considering it. The model rarely pushes back on the frame of the question itself. A good human advisor’s first response would have been: why are you even looking at that?
How to actually use it
A few things that measurably change the output.
Put every number in. Age, income, debts with their interest rates, account balances, monthly surplus, timeline. Without those, you’re running an expensive search engine.
Make it interrogate you. Add this line: “Before answering, ask me for any information you need that I haven’t provided.” That single sentence does more for output quality than any other prompting trick, and it’s the closest thing to a fix for not knowing what to ask. It converts the model from an answer machine into something resembling an intake conversation.
Demand the counterargument. “Under what conditions would this advice be wrong?” Models will argue against themselves readily when asked. They almost never volunteer it.
Verify the hard numbers. Contribution limits, bracket thresholds, phase-out ranges, and product terms move every year. Check them against the IRS or the institution directly. Models are strong on durable principles and weak on this year’s specifics.
The takeaway
The story here isn’t that AI financial advice is bad. It’s that the advice is genuinely good and only a narrow slice of people can extract it. The bottleneck moved off the tool and onto the user’s prior knowledge.
This almost certainly generalizes. Law, medicine, career strategy — anywhere expert advice matters, the same structure should hold. We built a technology that was supposed to democratize expertise, and it may be doing the opposite: handing the largest gains to people who needed the help least. Think about the last thing you asked a model. Did your question contain enough for it to actually help you?
Comments
Loading comments...