Benchmarking — Comparing performance or methods against a reference point, in four types by distance: internal, competitive, functional and generic.

Benchmarking: The Four Types, the Process Models Compared, and What Transfers

Robert C. Camp, Xerox 1979, published 1989 Moderate Complexity

Benchmarking is the comparison of an organization's performance or methods against a chosen reference point, in four types by distance: internal, competitive, functional and generic.

Before you start

Is this your framework?

Benchmarking answers one question: how do we compare, and what explains the difference. Both halves matter. A comparison without an explanation gives you a target and no route to it.

It is not a competitive strategy tool, though it is often mistaken for one. Learning how a rival runs its warehouse tells you nothing about whether the market is worth being in. The table below routes to what does.

Matching your actual problem to the right framework.
If your real problem is…You probably want
We need to know whether this market is structurally worth competing inPorter's Five Forces — industry attractiveness, not operational comparison
Compare Benchmarking and Porter's Five Forces
We need to see where rivals sit relative to us on the attributes buyers care aboutCompetitive Positioning Map — relative position on two axes, rather than process detail
Compare Benchmarking and Competitive Positioning Map
We already know the gap and need to close it in a specific processLean Six Sigma — a method for removing variation and waste once the target exists
We need to pick which measures to track in the first placeKey Performance Indicators — selecting and owning measures, which benchmarking then compares
We need to understand our own process before comparing it to anyoneValue Stream Mapping — you cannot benchmark a process you have not documented
We want to find where we are structurally slow, not where we differ from othersTheory of Constraints — the internal bottleneck, which no external comparison will name for you
We suspect we are behind and need evidence of by how much, and whyBenchmarking — you are in the right place

What Is It?

Pick something you do. Find somebody who does it better. Work out what they do differently, and whether any of it can be carried back. That is the whole method. The last part is where the difficulty lives.

The practice began at Xerox around 1979. The company found that Japanese rivals were selling machines for roughly what Xerox spent to build them. Taking rival products apart explained some of it. The rest was in how they made them. So Xerox began comparing processes rather than products, and then began comparing them against firms outside its own industry.

That move is the one worth understanding. The best available example of a process is usually not a competitor, because competitors face the same constraints you do and have converged on similar answers. Someone in an unrelated industry, solving a structurally similar problem under different pressures, is more likely to have arrived somewhere genuinely different.

The catch is symmetrical. The further afield you look, the more likely you are to learn something new. You are also less likely to be able to use it. Whatever made the practice work at the source is rarely visible from outside. Often it is not fully understood at the source either.

So there are two distinct activities under one word, and confusing them is the standard failure. Comparing numbers tells you a gap exists. Comparing methods tells you what causes it. Programs that do the first and skip the second produce a target, a slide, and a team told to improve by 30% with no account of how.

The four types of benchmarking drawn as nested scopes, with your own organization at the center: internal encloses it, competitive encloses internal, functional encloses competitive, and generic encloses them all
Moving outward, data gets easier to obtain and the legal exposure falls, while the practices get harder to carry back. The easiest comparison to run is rarely the one that transfers

Quick Reference

Complexity
Moderate (5/10)
Time to Decision
3-16 weeks
Data Required
High
Team Size
5-30
Objectivity
Medium-High
Learning Curve
2-4 weeks

The types

Two authorities, two different fours

Search for the types of benchmarking and you get two confident, different answers. Neither is wrong. They are cutting the same thing along different lines, and knowing which one someone is using saves an argument.

The two taxonomies, what each sorts on, and how they line up.
TaxonomyWhat it sorts onThe four
Camp and ASQ
the older lineage
Who you compare with, ordered by distance from your own organizationInternal, your own units. Competitive, direct rivals. Functional, the same job in other industries. Generic, any process anywhere.
APQC
the two-axis cut
Who and what, treated as separate questionsInternal and external for who. Performance and practice for what. Any study is one of each, so there are really four combinations rather than four kinds.

The two fit together once you see that the older list only answers the first question. Internal, competitive, functional and generic are four points on the who axis. Performance and practice cut across all four of them.

The two questions any benchmarking study answers, and what each is good for.
The cutWhat it producesWhat it cannot tell you
Performance
comparing numbers
Whether a gap exists and how big it is. Cheap, fast, often available from published sources or a paid study.Why. It will also mislead if the populations differ in ways the number hides.
Practice
comparing methods
What the other party does differently, and which parts of it depend on their context rather than their skill.How big your gap is. It needs the performance comparison to know whether the difference matters.
Internal
your own units
Fast, honest data, no legal exposure, and practices already proven in your context.Whether any of your units is good by outside standards. A best-of-five may be nobody's best.
External
anyone else
A reference point outside your own habits, which is the only way to discover you are collectively slow.Whether the comparison is like for like. This is where most benchmarking numbers go wrong.

Which pairing to run, and in what order

The cheap starting point is internal performance. Compare your own units, since that data is free and you can trust it. It tells you your spread. If the spread is wide, the gain is available without leaving the building, and an internal practice study will find it.

If the spread is narrow, everyone is doing the same thing and internal comparison has nothing left to give. That is when to go outside, and it is also when the like-for-like problem starts. Go external for performance to size the gap, then external for practice to explain it. Expect the second to take several times as long.

Methodology

The process models compared

The published process models differ mostly in how finely they slice the same sequence. All of them plan, collect, compare, and then try to get something adopted. Where they genuinely differ is how much weight they put on that last part, which is the part that fails.

The common process models, their shape, and what each emphasizes.
ModelShapeWhat it emphasizes, and the trade-off
Camp
Xerox, published 1989
Ten steps in four phases: planning, analysis, integration, actionThe most detailed. It is the only one that gives integration its own phase, where findings get accepted and turned into goals before anyone implements. Demanding to run in full, which is usually why people reach for a shorter model.
Spendolini
1992
Five stages, drawn as a cycleDeliberately simplified and generic, built to be adapted rather than followed exactly. Easier to start; offers less help at the point where findings meet resistance.
APQC
clearinghouse practice
Four phases: plan, collect, analyze, adaptBuilt around access to a comparison pool and a code of conduct governing the exchange. Strongest on how to obtain data legitimately; assumes you have somewhere to obtain it from.
Generic quality-cycle versionsMapped onto plan-do-check-actFamiliar to anyone already running continuous improvement, and folds benchmarking into an existing cadence. Loses the specific guidance on partner selection and comparability, which is where the errors are.

The choice matters less than the two steps everyone skips

Any of these models works. What separates studies that change something from studies that produce a slide is not which model was followed but whether two particular steps were done properly.

First, check comparability before collecting anything. If the two populations differ in a way the headline number hides, every later step compounds the error. The more rigorous the process, the more convincing the wrong answer looks. Second, treat adoption as work rather than as an outcome. Camp gives integration a phase of its own for a reason. The practice does not install itself, and the evidence on why is unusually clear.

Core Features

  • A defined subject: one process or measure, scoped narrowly enough to compare
  • A comparison partner: chosen deliberately, at a known distance from your own context
  • A comparability check: what differs between the populations before any number is trusted
  • A gap, quantified: the size of the difference, and whether it is still moving
  • A causal account: what the other party does differently, and what makes it work there
  • A transfer plan: which parts are portable, and who owns installing them

Worked example

A 34% gap that was mostly arithmetic

An illustrative composite. A grocery retailer in Cape Town, South Africa, with four distribution centers. Order fulfillment costs were rising and management suspected the picking operation was the cause. They started, as most do, with a published performance benchmark.

What each stage of the comparison produced, and what it cost.
StageWhat it showed
The performance benchmarkAn industry study put lines picked per hour at 148. Their own average was 98, a gap of 34%. The number was accurate and the conclusion drawn from it was wrong.
Chasing the numberA slotting and layout project ran for seven months at about R2.4 million. Picks reached 112. The gap narrowed and did not close, and nobody could say why the remaining distance existed.
The comparability check that had been skippedThe benchmark population had average order sizes roughly three times larger. Lines per hour rises with order size regardless of how well anyone picks. Normalized for order profile, the real gap was about 9%, not 34%.
The practice studyTwo site visits to a pharmaceutical distributor — a different industry with a similar order profile. The difference was not layout. It was wave release: batching orders into scheduled release windows rather than picking on arrival.
What transferred, and what did notWave release was implemented for about R310,000 in scheduling changes and training, taking picks to 129. The partner's automated replenishment triggers were also identified as superior and were not attempted, because they depended on a supplier data feed the retailer had no way to obtain.

The expensive project was aimed at a gap that was mostly arithmetic

R2.4 million bought fourteen picks an hour. R310,000 bought seventeen. The difference was not effort or execution. The second project was aimed at a cause and the first at a number. The performance benchmark was accurate and useless on its own. Two thirds of the gap it reported came from comparing different order profiles.

Note also which comparison actually helped. Not a competitor: a pharmaceutical distributor, in an unrelated industry, facing a structurally similar problem. And note what did not come back. The best practice they found was correctly identified, correctly understood, and still not transferable, because it rested on something the source had and the recipient could not get. That is the ordinary outcome, not the unlucky one.

When to Use

  • You suspect underperformance and need evidence of size before committing budget
  • An improvement target is being set and there is no basis for the number
  • Internal units vary widely and nobody knows which one is right
  • A process has been optimized locally for years with diminishing returns
  • An investment case needs an external reference point to be credible
  • Entering a new market or category and needing a sense of the going rate

When NOT to Use

  • The aim is a genuinely new approach rather than catching up to a known one
  • No comparable population exists, or the differences swamp the similarities
  • Your own process is undocumented, so there is nothing defined to compare
  • The comparison would require exchanging sensitive data with direct competitors
  • The gap is already known and the real problem is that nothing gets implemented
  • Speed matters more than evidence and a decision is needed this week

In practice

How benchmarking goes wrong

Most of these are versions of one error: treating a number as a comparison when the two populations were never alike.

The recurring failure modes and their remedies.
Failure modeWhat it looks likeWhat to do instead
Not like for likeA confident gap that dissolves once you find out what the benchmark population actually containsEstablish comparability before collecting. Ask what would make this number differ for reasons other than performance.
Numbers without methodsA target handed to a team with no account of what the leader does differentlyPair every performance comparison with a practice study, or do not set the target.
Copying the visible partThe practice is adopted and the result does not follow, because what made it work was not the visible partAsk what has to be true for this to work, then check whether those things are true here.
Competitive contact without limitsDirect data exchange with rivals covering prices, costs, margins or wagesUse published sources or a neutral intermediary. Follow a code of conduct, and take legal advice before direct contact.
Benchmarking to the medianThe target is set at industry average, which locks in being averageDecide first whether the goal is parity or advantage. They lead to different comparison partners.
Study finishes, nothing changesA thorough report, accurate findings, and no adoptionTreat adoption as its own phase with an owner, not as the natural consequence of a good report.
Only ever looking at competitorsEvery comparison inside the industry, so the whole industry's blind spots are invisibleRun at least one comparison against a structurally similar problem in an unrelated industry.

Sourced

Evidence, and how to cite it

The practice started at Xerox around 1979. The book came a decade later.

Xerox began comparing itself against Japanese rivals at the end of the 1970s. It had found that they were selling at close to Xerox's own cost. Robert Camp brought the approach to the logistics operation in 1981, and it became standard across the company soon after. His book, which set out the ten-step process, came in 1989 from ASQC Quality Press. Camp has said the formal process does not appear in Deming's work, or in anything before Xerox started. So the common framing of benchmarking as a branch of the quality movement has the lineage backwards.

Camp, R.C. (1989) Benchmarking: The Search for Industry Best Practices That Lead to Superior Performance. Milwaukee: ASQC Quality Press.

The famous Xerox example is not competitive benchmarking, which is what Xerox started with.

The case everyone cites is Xerox studying warehouse order picking at L.L. Bean, an outdoor goods retailer. That is generic benchmarking across unrelated industries, and it came after the competitive teardown work, not instead of it. The distinction matters because the story is usually told to illustrate "learn from competitors" when it illustrates the opposite. Xerox ran comparisons across a large number of areas through the 1980s, and won the Malcolm Baldrige National Quality Award in 1989.

Camp (1989); contemporaneous accounts of the Xerox benchmarking program.

Identifying a best practice is the easy half, and the reason is not motivation.

Gabriel Szulanski studied 122 best-practice transfers across eight companies, using 271 observations. He wanted to know why a practice that works in one place fails to take hold in another. The usual explanation blames resistance and not-invented-here. The data did not support it. The main barriers were about knowledge: whether the recipient could absorb it, whether anyone could say why the practice worked at all, and how hard the source and recipient found it to work together. Note that this was transfer within single firms, where access and motivation are easiest. Across firms it is harder, not easier.

Szulanski, G. (1996) ‘Exploring internal stickiness: impediments to the transfer of best practice within the firm’, Strategic Management Journal, 17(S2), pp. 27–43.

There is no controlled evidence that benchmarking programs improve performance.

The case rests on a few celebrated company histories, Xerox foremost among them, and those are picked on the outcome. Firms that ran benchmarking programs and did not recover do not get written up. Survey research on adoption exists. It cannot separate the effect of benchmarking from the effect of being the kind of firm that runs structured improvement programs. Treat it as a disciplined way of finding out where you stand, which it is, rather than as a proven intervention.

Assessment of the management literature as of 2026; the widely cited productivity-gain figures by benchmarking type circulate without traceable primary sources.

Competitive benchmarking has legal limits that the process models mostly gloss.

Sharing sensitive information with rivals can attract antitrust attention whatever the reason for it. Prices, costs, margins, wages and the division of markets are the areas to be most careful about. This is why benchmarking codes of conduct exist. It is also why competitive comparisons usually run through published filings, third-party studies, or a neutral party that pools and anonymizes the data. This page describes the practice and is not legal advice. Direct contact with a competitor is worth clearing with counsel first.

APQC benchmarking code of conduct; general competition law principles.

How to cite it.

Harvard: Camp, R.C. (1989) Benchmarking: The Search for Industry Best Practices That Lead to Superior Performance. Milwaukee: ASQC Quality Press.
APA: Camp, R. C. (1989). Benchmarking: The search for industry best practices that lead to superior performance. ASQC Quality Press.
For the types, name which taxonomy you mean and cite ASQ or APQC accordingly, because they differ. For transfer failure, cite Szulanski (1996) in the Strategic Management Journal.

Key Strengths

  • Replaces opinion with a number: an external reference beats internal argument about what is possible
  • Finds practices you would not invent: particularly from outside your own industry
  • Sizes the prize before you spend: the gap tells you whether the work is worth doing
  • Internal comparison is nearly free: the data is yours and already trustworthy
  • Travels well to investment cases: an outside reference point is credible to people who were not in the room

Key Weaknesses

  • Comparability is fragile: the number is only as good as the like-for-like check nobody does
  • Aims at parity: catching up to a known practice is not the same as getting ahead
  • Best practice rarely transfers whole: and the reasons are structural, not attitudinal
  • Competitive data is guarded and constrained: legally as well as practically
  • Slow where it is useful: practice studies take months, which is why programs drift to numbers only

Sequencing

What to run before and after

Benchmarking sits between knowing what you do and changing it. Run before the first and it compares nothing; run without the second and it produces a report.

Before

Document your own process and decide what you are measuring

A process you have not mapped cannot be compared, and a measure you have not defined cannot be matched to somebody else's. Most comparability failures are really definition failures found late.

During

Choose partners deliberately, and check the comparison holds

Partner choice determines what you can learn and how much of it will transfer. The nearer the partner, the more portable the practice and the less new it is.

After

Treat installing the practice as its own project

This is the step with the clearest evidence against it succeeding by default. Give it an owner, a plan, and an explicit account of what has to be true locally for the practice to work.

Common questions

Benchmarking: quick answers

What are the four types of benchmarking?

It depends which authority you follow, and the two common answers differ. The Camp lineage, carried by ASQ, sorts by who you compare with: internal, competitive, functional and generic. APQC sorts on two axes instead: internal and external for who, performance and practice for what. The second set is not a rival list so much as a different cut, and the table on this page lines them up.

What is the difference between performance and practice benchmarking?

Performance benchmarking compares numbers and tells you whether a gap exists and how large it is. Practice benchmarking compares methods and tells you what is causing the gap. Running the first alone is the commonest mistake: you get a target with no route to it, and teams end up guessing at what the leader does differently.

What is competitive benchmarking?

Comparing your performance or methods directly against named rivals in your own market. It produces the most immediately relevant comparison and is the hardest to do, because the data is guarded and because exchanging certain information with competitors carries antitrust exposure. Most organizations get competitive comparisons from published filings, third-party studies or a neutral intermediary rather than directly.

Is competitive benchmarking legal?

Comparing publicly available information is routine. Direct exchanges with competitors are where the risk sits, particularly anything touching prices, costs, margins, wages or market allocation, which can attract antitrust scrutiny regardless of intent. APQC publishes a benchmarking code of conduct covering this ground. Anything involving direct contact with a competitor is worth running past counsel first; this page is not legal advice.

What is Camp's benchmarking process?

Ten steps grouped into four phases: planning, analysis, integration and action. Planning decides what to benchmark and against whom. Analysis measures the gap and projects where it will be in future. Integration gets the findings accepted and turned into goals. Action implements, monitors and recalibrates. It is the most detailed of the common models and the most demanding to run.

What is the difference between a benchmark and benchmarking?

A benchmark is a reference number, such as an industry median cost per unit. Benchmarking is the process of finding a comparison, measuring the gap and working out what explains it. Buying a benchmark report gives you the first and none of the second, which is why so many benchmarking programs stop at a slide showing a gap.

How long does a benchmarking study take?

A performance comparison against published data can be done in two or three weeks. A practice study involving site visits, interviews and process mapping usually runs two to four months, and Camp's full ten-step cycle takes longer still. The gap between those numbers is why many programs quietly become performance-only exercises.

How do you cite benchmarking sources?

For the process and the Xerox origin, cite Camp, R.C. (1989) Benchmarking: The Search for Industry Best Practices That Lead to Superior Performance, ASQC Quality Press. For the four types, say which taxonomy you mean and cite ASQ or APQC accordingly. For evidence on why identified best practices fail to transfer, cite Szulanski (1996) in the Strategic Management Journal.

Deep Resources