The Briefing · Issue 5 · September 1, 2026 · 7 min read
Same Prompt, Four Models, Four Very Different Answers
I gave ChatGPT, Claude, Gemini, and Copilot the identical research assignment at their highest settings. The differences should change how you pick your tools.
Here’s a pattern I see in almost every bank I work with. Someone picked an AI tool a while back. Usually it was whichever one showed up first, or whichever one came bundled with something the bank already paid for. They use it on the default setting, they get decent answers, and they stop there. They never test the alternatives. They never switch on the top tier. They have no idea what they’re not seeing.
So I ran the test for you.
I gave four models the exact same assignment, word for word, each one set to its highest performance level: ChatGPT, Claude, Gemini, and Microsoft Copilot. These top settings go by different names, things like deep research and extended thinking, but the idea is the same. You trade a fast answer for a researched one, and most users have never clicked the button.
The assignment was a real one, and it has a twist I enjoyed. I asked each model to investigate how Microsoft Copilot treats a bank’s data. Yes, that means I asked Copilot to investigate itself.
"I need a deep dive on how Copilot treats my bank's data and private info. It should be highly detailed with sources cited so I can dig in to make the right usage and risk assessments. The output should be a Word document."
One prompt. Four flagship models. Here’s what came back.
Where all four agreed
Before the differences, the consensus, because it matters. All four reports landed on the same core facts. Microsoft does not use your prompts, responses, or tenant data to train its foundation models. Copilot only sees what each employee already has permission to see, which means your real exposure is years of sloppy SharePoint sharing, not Microsoft’s data handling. Web searches leave the protected boundary and run under different legal terms. And your prompts are stored, retained, and discoverable, so they’re bank records.
When four independent tools reach the same conclusion from the same question, you can take that consensus to your risk committee with confidence. That alone is a reason to use more than one model on anything important.
Where they split
The consulting binder. The longest report by far, and the most operational. It produced a full risk register with inherent and target ratings, a data classification decision matrix, a 90 day implementation roadmap, detailed test scripts, starter policy language, and a RACI chart assigning every control an owner. It reads like a deliverable you'd pay a consulting firm real money for. The gaps: it never mentioned the documented security incidents, and its 59 sources are named but not linked, which makes verification slower than it should be.
The community banker's briefing. The only report written specifically for community banks and credit unions, in plain second person, organized around what your examiner will ask. It was also the only one to report the negative history: a zero click vulnerability called EchoLeak that was found and patched in 2025, and a separate flaw where Copilot accessed files without writing audit log entries, which Microsoft fixed without telling customers. It even included a section flagging its own weakest claims so you know what to verify. 72 sources, mostly primary, with links.
The engineer's whitepaper. The shortest and the most deeply technical. It was the only report to cover the encryption tradeoff that matters at the architecture level: Double Key Encryption locks Copilot out of your files completely, while Customer Managed Keys keep AI working, so your key strategy decides what AI can touch. Two cautions. It wrote for a Wall Street audience, citing broker-dealer rules and European regulations a community bank doesn't face. And its sourcing was the weakest of the four, leaning on help desk pages, reseller blogs, and a Reddit thread.
The self assessment. Copilot reviewing Copilot was better than I expected. Its framing is genuinely useful: Copilot is an acceleration layer over your existing permissions, so strong controls get more useful and weak ones get easier to exploit. Then the part that stopped me. Because Copilot runs inside the tenant, it found and cited an internal draft AI policy document in its own report. No other model could do that. But every one of its sources was Microsoft's own documentation, and like most of the field, it said nothing about the security incidents.
What this teaches you
Notice what just happened. The same question, asked the same way, produced a consulting engagement, an examiner prep briefing, an architecture whitepaper, and a vendor self review. None of them is wrong. They’re different products, built by different companies with different instincts about who is asking and what good looks like.
That has three practical consequences for your bank.
Match the model to the job, not the habit
Need an implementation plan with owners and dates? That 40 page structure is your friend. Need to brief your board or prep for an exam? You want the one that writes for your audience and shows you the ugly parts. Need to settle a technical architecture question? The deep technical treatment wins. The right answer changes with the task, and you only learn each model's personality by trying it.
Use a second model on anything high stakes
Three of the four reports never mentioned that a serious vulnerability had been found and patched. If you'd read any one of those three alone, you wouldn't know. One model gives you an answer. Two models give you an answer plus a view of what the first one left out. For a risk assessment, a vendor decision, or anything going in front of your board, that second read is cheap insurance.
Turn on the top tier before you judge
Every one of these reports came from the highest setting, and none of them resembles what the same tools produce on a quick default chat. If your opinion of AI was formed on a free tier answering in eight seconds, you haven't actually evaluated these tools. Run one real task through the research tier before you decide what your bank's tools can and can't do.
The default setting is a different product from the top setting. Most people have only ever met the default.
Read all four reports
Don’t take my word for any of this. Here are the four reports exactly as the models delivered them, converted to PDF and otherwise untouched. Skim all four side by side and you’ll feel the differences within a few pages. Consider sharing them with your IT and compliance teams, because the Copilot findings themselves are worth their time.
One more thing. If your team has been living on a single model’s default setting, this is exactly the kind of side by side I run in my working sessions with banks, on your tasks instead of mine. It changes how people pick their tools in about an hour.
Click here to get insight delivered directly to your inbox →
Ben