Back to Blog

Asking Copilot a Data Question in Power BI - What It Does Well and What Trips It Up

September 10, 20267 min readMichael Ridland

The dream with business intelligence has always been the same. A manager who is not a data analyst types "how did the Queensland region do last quarter compared to the year before" and gets a straight answer with a chart, without waiting three days for someone in the data team to build a report. Every BI vendor has promised some version of this for a decade, and most of the attempts were disappointing enough that people stopped believing it. Copilot in Power BI is the closest I have seen it get to actually working, and I want to give you a straight read on it, because the gap between the demo and the reality is where teams get burned.

The feature itself is simple to describe. You ask Copilot a question about your data in plain English, and it responds with an answer, usually a visual and some narrative, drawn from your semantic model. The Microsoft documentation walks through the mechanics. What I want to talk about is whether it delivers on the promise, because I have now watched a fair few Australian teams switch it on, and the results split cleanly into two camps.

What it does genuinely well

When it works, it is properly good. A well-formed question against a clean model comes back with a sensible answer in seconds. "Show me sales by product category for the last six months" gets you the right chart without anyone touching the report builder. For the everyday questions that make up the bulk of what business users actually want, the ones that would otherwise generate a request to the data team or a fiddly self-serve attempt, it removes the friction almost entirely.

The narrative summaries are the underrated part. Copilot does not just draw the chart, it tells you what the chart says. "Sales grew twelve per cent, driven mainly by the Homewares category, while Outdoor declined." For someone who finds a chart intimidating, having the takeaway written out in words is the thing that makes the data usable to them at all. I have seen this genuinely change who engages with reports in an organisation, pulling in people who had quietly opted out of self-serve BI because they never felt confident reading the visuals.

The follow-up capability matters too. You can ask a question, then narrow it, then narrow it again, and it holds the thread. That conversational back and forth is how people actually explore data when they are thinking out loud, and it is a much more natural fit than the old model of building a fresh visual for every question.

What trips it up, and it is always the same thing

Here is the honest bit. When Copilot gives a poor answer in Power BI, the cause is almost never Copilot. It is the semantic model underneath. And this is the single most important thing to understand before you roll it out, so I will be blunt about it.

Copilot can only be as good as the model it is asking questions of. If your measures are named things like "Measure 1" and "Calc_v2_final", Copilot has no idea what they mean and neither would a human. If you have three different fields that all sort of represent revenue and no clear indication which is the real one, Copilot will pick one, and it might pick wrong, and it will present the wrong answer with exactly the same confidence as a right one. That last part is the dangerous bit. A human analyst who is unsure will hedge. Copilot states its answer plainly whether it is right or not.

So the model has to be built for humans and machines to understand. Clear field names. Well-defined measures with obvious meanings. Descriptions on the tables and columns so the intent is explicit. A single unambiguous source of truth for each concept that matters. When we get called in because "Copilot is giving wrong answers", the fix is essentially always in the model, not in Copilot. This is where a lot of our Power BI consulting work has shifted, because getting a model ready for natural language questioning is a real discipline and most models built over the years were not built with it in mind.

The trust problem you have to plan for

There is a deeper issue than wrong answers, and it is about trust. When a person asks a question and gets a confident answer with a nice chart, they believe it. That is human nature. The presentation carries authority whether or not the number behind it is correct. So if your model has a subtle flaw, or Copilot picks the wrong field for an ambiguous question, you now have a business user making a decision on a wrong number they have no reason to doubt.

This is why I tell clients that turning Copilot loose on a shaky model is worse than not having it at all. At least with the old model, a wrong report went through a data person who might have caught the error. Natural language questioning removes that checkpoint and puts confident answers directly in front of decision makers. The technology is not the risk. The absence of a clean, governed model underneath it is the risk, and the technology just amplifies whatever state your data is in.

The organisations that get value from this are the ones that treat the model as the product. They invest in getting it clean, well described and trustworthy first, and then Copilot becomes a genuine multiplier on that investment. The ones that switch it on hoping it will paper over a messy model get burned, sometimes quietly, which is the worst way to get burned because you do not find out until a decision has already gone wrong.

Licensing and the practical realities

A couple of practical notes. This capability sits behind the paid Power BI and Fabric capacities, so it is not a free add-on and you should factor that into any rollout plan. Check what your current licensing actually gives you before you promise the business a feature it cannot access yet.

The quality also varies with the shape of your data. Straightforward questions against a well-built star schema work beautifully. Convoluted questions, or questions that require reasoning across a tangled model with unclear relationships, are where you see the wheels wobble. Set expectations accordingly. This is a strong tool for the common case, not a magic analyst that handles every edge case a human expert would.

How to actually roll it out

If you want this to land well, do it in an order. Pick one important semantic model, the one that answers questions the business genuinely cares about. Get that model properly clean: sensible names, clear measures, good descriptions, one source of truth per concept. Then enable Copilot on just that model, with a small group of real users, and watch what happens. See which questions it handles and which it fumbles, and use the fumbles to keep improving the model.

That contained approach beats switching Copilot on across the whole tenant and hoping. A big-bang rollout across a pile of inconsistent models is how you generate a wave of wrong answers and a durable loss of trust that is very hard to win back. Start narrow, get it genuinely good, then expand. We often pair this with a bit of AI training for the people who will use it, because knowing when to trust an answer and when to sanity-check it is a skill in itself, and it is one worth teaching deliberately.

Natural language questioning of your data is finally good enough to be worth taking seriously, which is not something I would have said a couple of years ago. But the value is entirely gated on the quality of what sits underneath it. Get the model right and it is a real step change in who can use your data. Skip that and it is a confident-sounding liability.

If you want help getting your Power BI models ready for this, or you have switched Copilot on and the answers are not landing, that is exactly the kind of thing we sort out. Have a look at our business AI and data services or get in touch and we will work through it with you.