What should be defined
- The business question and why mystery shopping is the appropriate method.
- Standards, scenarios and behaviors that can be observed objectively.
- The sample of locations, channels, times, profiles and repetitions.
- Evaluator and instrument preparation, calibration and quality control.
- Responsible data use, analysis, action and new measurement.
1. Start with the decision, not the questionnaire
Describe which decision will change because of the evidence: correct a reception process, compare execution across channels, validate a new service promise or prioritize training. Turn the problem into bounded questions. “Do they provide good service?” is ambiguous; “Does the team confirm the need before recommending an option?” is observable.
Review existing complaints, timing, sales, returns, surveys, calls and standards. Mystery shopping helps observe a designed experience under specific conditions; it does not estimate satisfaction, market share or the cause of a behavior by itself. Combine it with other methods where the decision requires them.
- Decision, owner and date when it must be made.
- Behavior, condition or outcome that needs observation.
- Relevant population, channel and moment.
- Existing data and the gap new evidence will cover.
- Limits of what one visit can conclude.
Choosing a method before framing the question produces precise data about something that may not matter.
2. Separate observation, perception and explanation
Mystery shopping records what occurred under a scenario. A survey captures what someone reports, an interview explores meaning, operational data shows outcomes, and open observation describes context. No source contains the whole experience and each introduces different bias.
Build a question-to-method matrix. If the objective is to know whether a warranty is offered, observation may suffice. To understand why an offer creates distrust, add interviews. To estimate how frequent an opinion is within a population, design an appropriate survey and sample.
- Question and evidence needed to answer it.
- Available primary or secondary source.
- Most direct method and corroborating alternative.
- Bias, coverage and cost of each option.
- Rule for integrating apparently contradictory results.
3. Translate standards into observable scenarios and criteria
A standard such as “be friendly” depends too heavily on interpretation. Describe behavior: greeting within an interval, confirming a need, explaining conditions, checking understanding or agreeing a next step. Avoid exact scripts where several behaviors can meet the intent.
Design realistic, ethical and comparable scenarios with profile, need, objection, limits and planned ending. The rubric separates fact, evidence and comment. Include “not applicable,” “not observable” and context; forcing a score invents precision. Test that each criterion is operationally controllable and improvable.
- Standard and its reason for the customer or business.
- Observable behavior and acceptable evidence.
- Scenario, question, objection and permitted ending.
- Scale with clear anchors and non-applicable states.
- Factors outside the control of the person or location.
A long rubric is not rigorous if two reasonable evaluators interpret it differently.
4. Design a sample representing the defined operation
List relevant locations, channels, days, periods, request types and profiles. A convenience sample may concentrate on accessible areas or quiet hours and create a misleading picture. Distribute visits according to the question and retain randomness where feasible.
Define the unit of analysis: visit, location, channel, shift or period. Several visits to one branch are not automatically independent if they share a shift, campaign or incident. Report effective count, missing coverage and substitutions. Do not generalize to all of Panama when the sample covers selected locations.
- Frame of locations, channels, schedules and profiles.
- Relevant strata and planned count for each group.
- Selection, substitution, repetition and exclusion.
- Seasonal events, campaigns and operational changes.
- Achieved coverage and population the result actually represents.
5. Prepare evaluators and the instrument through a pilot
Recruit evaluators able to represent the scenario and record facts without unnecessarily altering the interaction. Explain purpose, permitted behavior, safety, expenses, evidence, privacy and exit from unexpected situations. Avoid incentives rewarding fault-finding or extreme scores.
Run a pilot to find ambiguous criteria, unrealistic timing and missing fields. Compare how several people score the same example and correct instructions. During fieldwork, review duration, location, consistency, improbable patterns and evidence; validate exceptions without revealing unnecessary information to operations.
- Evaluator profile, conflicts and preparation.
- Scenario, expenses, purchases, returns and limits.
- Scoring examples and practice in factual evidence.
- Pilot, corrections and final instrument version.
- Field controls and procedure for invalidating a visit.
6. Design ethics, privacy and data protection from the start
The ICC/ESOMAR Code establishes duty of care, minimization, privacy, transparency, accountability and fitness for purpose. In Panama, ANTAI explains obligations and rights for personal-data processing under Law 81 of 2019 and its regulation. Define which data is actually necessary, who accesses it, how long it remains and which uses are allowed.
Never assume audio, images, names or individual identifiers. Assess necessity, applicable basis, permissions, security and employment effects with qualified advice. Keep research results separate from incompatible uses and avoid publishing fragments that identify a person through the combination of branch, time and description.
- Purpose, responsible organization, data and affected people.
- Minimization, access, transfer, retention and disposal.
- Permission and review for audio, images or location.
- Handling of minors, health and other sensitive contexts.
- Aggregate use, anonymization and boundaries for employment decisions.
This framework supports responsible design; the organization must validate its concrete legal and employment obligations.
7. Analyze patterns without hiding variation or uncertainty
Define calculations, segments and comparisons before fieldwork. Do not change weights to produce a more comfortable story. Show base, observation count, dispersion, non-applicable fields and coverage differences. An average may hide that a process works in one channel and fails in another.
Separate observed compliance, evaluator comment, evidence and proposed explanation. Corroborate with surveys, operations or interviews before assigning cause. AAPOR recommends controls throughout the survey lifecycle; the same traceability principle helps uncover programming, collection, coding and analysis errors in mixed studies.
- Base and period for every indicator.
- Results by criterion and planned segments.
- Variation, missing values and invalidated visits.
- Redacted examples that illustrate without replacing the pattern.
- Conclusion, alternative and limitation for each recommendation.
8. Turn the report into learning and new measurement
Prioritize a few gaps by impact, frequency and controllability. Identify whether the likely cause sits in process design, capacity, inventory, information, tools, leadership or skill. Do not use training to fix a system that demands behaviors incompatible with available time or resources.
Assign an owner, date, change and evidence. Share results with context, recognize strengths and allow teams to explain conditions. Measure again with comparable instruments after a reasonable period, recording any sample or standard change. The goal is not to improve the score through rehearsal but the underlying experience.
- Prioritized gap, impact and cause to validate.
- Action on process, tool, capacity or skill.
- Owner, date and implementation evidence.
- Communication and support for teams and locations.
- New wave, comparability and improvement criterion.
Frequently asked questions
Questions that should be settled before acting
How many visits does a program need?
It depends on the locations, channels, times and scenarios to represent, expected variation and the decision. A fixed number without a sample frame does not guarantee usefulness.
Should visits be a surprise?
The exact date may remain undisclosed to observe usual operations, but the organization should establish the program, authorization, safeguards and result uses according to its context and obligations.
Can branches be ranked?
Only where sample, scenarios, timing and coverage are sufficiently comparable. Even then, show uncertainty and context; a ranking can exaggerate small differences.
Which evidence should an evaluator attach?
The minimum needed to validate date, scenario and permitted facts: a receipt, structured record or authorized material. Do not collect personal data or recordings by habit.
How often should the program repeat?
After enough time to implement changes or when risk and operations justify it. Measuring too early rewards temporary preparation; measuring too late delays learning.



