Skip to content

Sample deliverable · September 2026

Power BI Copilot readiness report

Northstar Revenue Analysis · Priority semantic model · Executive commercial users · 20 representative questions

Illustrative sample based on a synthetic semantic model. Not a client result.

Executive verdict

38/100

Not ready for consequential use

Decision

Do not release the priority revenue use case to the executive group yet. The model can produce useful exploratory responses, but it cannot consistently prove which revenue definition, calendar basis, or permission path produced the answer.

Fastest credible path: resolve the two business-policy ambiguities, constrain the AI-visible surface, replace implicit calculations, and establish a 20-question regression gate before a controlled pilot.

Readiness heatmap

Seven dimensions, one release decision.

Scores summarize evidence strength. Critical security or evaluation gaps override a numerically higher readiness band.

Model structure68Controlled
Business terminology52Material gaps
Measures & foundations31High risk
AI context & instructions44Material gaps
Verified answers38High risk
Security & governance22Critical
Evaluation & lifecycle11Critical

Synthetic model context

What was assessed

Business use case
Explain revenue performance and variance for executive commercial reviews
Model
Northstar Revenue Analysis v12 (synthetic)
Target users
12 executive and finance consumers across three permission roles
Question set
20 representative prompts covering revenue, variance, calendar, region, and product
Evidence reviewed
Model metadata, measure definitions, instructions, AI schema, permission matrix, and test records
Out of scope
Capacity sizing, tenant configuration, report design, and production implementation

Prioritized findings

Every finding connects evidence to business risk and action.

F-01

Critical

Two competing revenue measures can both answer “revenue.”

Evidence

The synthetic model exposes Net Revenue and Reported Revenue with overlapping descriptions. Five representative prompts select different measures without an explicit basis.

Business risk

An executive comparison can change when phrasing changes even though both calculations are technically valid.

Recommended action

Name one default revenue measure for this use case; document exclusions, grain, calendar, filters, and owner; constrain alternates to explicit contexts.

F-02

Critical

Calendar basis is unresolved for year-to-date questions.

Evidence

The Date table contains fiscal and commercial calendars. Instructions say “use fiscal dates,” while the priority report defaults to commercial periods.

Business risk

A correct aggregation can answer the wrong time-basis question and materially change the result.

Recommended action

Resolve the policy with the metric owner, encode the default in the measure layer, and add explicit calendar variants to the question set.

F-03

High

Implicit measures bypass approved business logic.

Evidence

Amount and Units remain summarizable columns. Generated visuals can sum them directly instead of using governed measures.

Business risk

Copilot can produce a plausible total that omits exclusions, currency logic, or grain safeguards embedded elsewhere.

Recommended action

Hide raw aggregatable columns from the business surface and replace priority calculations with explicit measures.

F-04

High

The AI data schema is too broad for the chosen question set.

Evidence

The schema exposes 14 tables and 186 fields; the 20 priority questions require 5 tables, 24 fields, and 9 measures.

Business risk

More plausible candidates increase field-selection ambiguity and make failures harder to diagnose.

Recommended action

Create a question-driven AI schema and test that included fields work and excluded fields are not selected.

F-05

High

AI instructions conflict with model definitions.

Evidence

One instruction defines “active account” using a 30-day window; the approved measure and glossary use 90 days.

Business risk

The instruction layer can steer interpretation away from the governed calculation and obscure the source of disagreement.

Recommended action

Remove calculation logic from instructions, point to the approved measure, and add a conflict-review step to model changes.

F-06

Critical

No repeatable golden-question regression exists.

Evidence

Testing consists of ad hoc screenshots. Expected values, required filters, tolerances, evidence, and release criteria are not versioned.

Business risk

A model or instruction change can silently improve one prompt while breaking another consequential question.

Recommended action

Create a 20-question evaluation set with expected answer contracts, failure categories, and an owner for disposition.

F-07

Critical

Permission behavior has not been tested through the AI path.

Evidence

Report-viewer roles were tested, but verified-answer and conversational experiences were not exercised for restricted-region users.

Business risk

A report permission test does not establish that every AI feature and delivery channel behaves safely for each role.

Recommended action

Run allow, deny, and edge-case questions using representative identities in the exact intended consumption channel.

Example answer contract

Define the answer before evaluating the AI.

Synthetic answer contract for net revenue year to date
FieldApproved contractEvidence owner
Business questionWhat is net revenue year to date, and how does it compare with prior year?Commercial finance
Measure[Net Revenue] and [Net Revenue PY]Semantic model owner
Source and grainCertified sales fact at transaction-day-product-customer grainData product owner
CalendarCommercial 4-4-5 calendar; completed periods onlyFinance policy owner
FiltersPosted transactions; approved revenue classes; user-permitted regionsMetric owner
ToleranceExact total to displayed precision; correct measure, calendar, and filters requiredEvaluation owner

Failed golden question · GQ-04

“What is revenue YTD versus last year?”

Observed
Selected Reported Revenue, fiscal YTD, and all transaction statuses.
Expected
Net Revenue, completed commercial periods, posted transactions, and prior-year comparison.
Failure class
Wrong measure + wrong calendar + missing filter. A numerically plausible answer is not accepted.

Evaluation rule

Score the answer path, not just the final number.

A matching value can be accidental or unstable. The evaluator records interpretation, measure, filters, calendar, security behavior, evidence, and repeatability. Any critical-contract miss fails the question even when the displayed total matches.

Evaluation excerpt

Five questions, five different failure modes.

Excerpt from synthetic golden question evaluation
IDQuestionResultPrimary failureEvidence
GQ-01Net revenue last completed weekPass—Correct measure, calendar, filter, value
GQ-04Revenue YTD vs prior yearFailMeasure + calendarPlausible value, wrong contract
GQ-07Top declining regionFailSecurityRestricted region appeared in verified visual
GQ-11Active accounts by segmentFailInstruction conflict30-day interpretation instead of governed 90-day measure
GQ-16Explain gross-to-net changeReviewScopeDescriptive answer lacked approved adjustment detail

Before and after

Before · broad instruction

Use revenue when users ask about sales. Prefer fiscal dates. Active accounts are accounts with recent sales.

Three ambiguous terms, no authoritative measure, and no precedence when the model disagrees.

After · governed context

For executive performance questions, use [Net Revenue]. Use completed Commercial Periods unless the user explicitly asks for Fiscal Calendar. Use [Active Accounts 90D]. If calendar basis is omitted and the question is consequential, state the basis in the response.

Business logic remains in explicit measures; instructions clarify selection and response behavior.

Target architecture

A controlled answer path with inspectable evidence.

01

Business question

Target user phrasing + decision context

02

AI context

Focused schema + instructions + verified-answer routing

03

Semantic contract

Measures + definitions + grain + calendar + security

04

Evaluation evidence

Expected path + observed result + release decision

30 / 60 / 90-day roadmap

First 30 days

Resolve and constrain

  • Name metric and calendar policy owners
  • Approve revenue and YTD answer contracts
  • Hide implicit aggregations
  • Reduce AI schema to priority surface

Days 31–60

Test and govern

  • Reconcile AI instructions with model logic
  • Configure bounded verified-answer candidates
  • Run role-based permission scenarios
  • Version the 20-question evaluation set

Days 61–90

Pilot and decide

  • Run controlled pilot with target users
  • Triage failures by model, context, policy, or product
  • Set release thresholds and review cadence
  • Decide Copilot, data agent, or custom-agent path

Assumptions and limitations

What this verdict does not claim.

  • All model names, values, user roles, questions, findings, and scores in this sample are synthetic.
  • The assessment is a point-in-time evidence review, not certification and not a guarantee of generated-answer accuracy.
  • AI output is nondeterministic. Passing a question set does not prove every possible prompt will produce an acceptable response.
  • Current Microsoft preview features, permissions, licensing, and capability behavior require revalidation before implementation.
  • Business-policy decisions remain with the accountable client owners; Refinity can expose ambiguity but should not invent policy.
  • Production configuration, data movement, privacy, legal, and broader platform architecture are outside this illustrative scope.

Next step

Test your own readiness before sharing any files.

Use the ungated scorecard for a directional view, or bring one consequential question to a technical review.