WHAT COUNTING COSTS
Four waves of measurement, and the same thing missing from all of them
It traces impact assessment from output counting to complexity-aware evaluation, and then asks the question none of the four waves asked: who is doing the measuring.
Progressive Impact Assessment · The History of Impact — A Developmental Story · 4361 words · 20 minutes
The counting impulse
There is something profoundly human about the desire to measure. Long before the first formal evaluation framework was codified, human beings were counting harvests, tallying trades, and assessing the outcomes of their collective efforts. The impulse to quantify is not, in itself, problematic. It is an expression of the longing to understand, to improve, to steward resources wisely. But somewhere along the winding path from ancient tallying sticks to modern impact dashboards, something essential was lost. The field became so good at counting that it forgot to ask what was worth counting in the first place.
This chapter tells the story of how the field arrived here — how impact assessment evolved from simple output counting to sophisticated outcome measurement, from narrow accountability frameworks to the emerging recognition that the things that matter most in human flourishing often resist the very instruments built to capture them. It is a developmental story in the fullest sense: not merely a chronological history, but a narrative of expanding awareness, each stage of the field's evolution representing both a genuine advance and a new set of blind spots.
Understanding this history is not an academic exercise. It is essential preparation for the work ahead, because the measurement frameworks a practitioner inherits shape not only how programs and interventions are evaluated. They shape how impact itself is conceived. And if the conception of impact is too narrow, too mechanistic, too divorced from the living complexity of human experience, then even the best-intentioned efforts will miss the mark.
The first wave: counting outputs
The earliest formal approaches to impact assessment were, by necessity, simple. In the mid-twentieth century, as governments and philanthropic organizations began investing heavily in social programs — public health campaigns, educational initiatives, poverty reduction efforts — the primary question was straightforward. Did we do what we said we would do?
This output-oriented approach asked: how many vaccinations were administered? How many students enrolled? How many meals served? These are not trivial questions. Outputs matter. If a program promises to deliver clean water to a village and fails to install a single well, no amount of philosophical sophistication can rescue it from that failure.
But output counting, as a comprehensive approach to assessment, has a fundamental limitation: it confuses activity with achievement. A literacy program that distributes ten thousand books has produced an impressive output. But if those books sit unread on shelves — if the program failed to account for the cultural context, the reading levels of the recipients, or the absence of lighting in their homes — then the output says almost nothing about the program's actual contribution to human flourishing.
The output era, roughly spanning the 1950s through the 1970s, was characterized by what might be called first-order measurement: the belief that tracking what was produced and delivered would give a sufficient picture of impact. This approach was often embedded in bureaucratic structures that prioritized accountability over learning. Program managers were rewarded for hitting output targets, and the incentive structure quietly discouraged anyone from asking whether those outputs were actually producing the changes they were designed to create.
The shadow of this first wave persists. Even today, many organizations, particularly those under intense funding pressure, default to output metrics because they are easy to collect, easy to report, and easy to understand. The seduction of the countable is powerful. But as the field matured, a growing chorus of voices began to insist that counting outputs was not enough.
The second wave: measuring outcomes
The shift from outputs to outcomes represented a genuine developmental leap. Beginning in the late 1970s and accelerating through the 1990s, evaluators and funders began asking a more penetrating question. What actually changed as a result of the intervention?
This was a crucial evolution. Instead of merely counting the number of students who attended a tutoring program, outcome measurement asked whether those students actually learned to read, whether their grades improved, whether they stayed in school. Instead of tracking the number of counseling sessions delivered, it asked whether participants reported reduced symptoms of depression, whether their relationships improved, whether they experienced a greater sense of agency in their own lives.
The outcome movement was fueled by several converging forces. The rise of program evaluation as a professional discipline brought methodological rigor to the assessment of social interventions. The work of pioneers like Michael Scriven, who distinguished between formative evaluation, aimed at improvement, and summative evaluation, aimed at judgment, gave the field a more nuanced vocabulary. Meanwhile, the growing influence of economists and policy analysts introduced cost-effectiveness analysis and cost-benefit analysis, demanding that social programs demonstrate not only that they produced outcomes, but that they produced outcomes efficiently.
Perhaps the most influential development of this era was the emergence of logic models, visual frameworks that map the causal chain from inputs, the resources invested, through activities, what the program does, to outputs, what is produced, then outcomes, the short and medium-term changes, and ultimately impacts, the long-term systemic changes. The logic model became the lingua franca of program evaluation, and for good reason: it forced program designers to articulate their theory of how change happens.
Yet the outcome era introduced its own shadows. The emphasis on measurable outcomes created a powerful incentive to focus on changes that could be captured by standardized instruments — survey scores, test results, behavioral indicators. Outcomes that resisted quantification — shifts in meaning-making, deepening of relational capacity, the slow emergence of collective wisdom — were systematically marginalized. Not because anyone believed they were unimportant, but because the measurement tools of the era could not see them.
There was also a subtle but significant epistemological assumption embedded in outcome measurement: the belief that causation could be cleanly established. If students who attended a program improved their reading scores, the program caused the improvement. This assumption drove the field toward increasingly rigorous research designs, most notably the randomized controlled trial, borrowed from medical research, which became the gold standard of evidence-based practice.
The randomized controlled trial is a powerful tool. In contexts where a discrete intervention produces discrete, measurable effects in a relatively short timeframe, it can provide robust evidence of causal impact. But human development is not a pharmaceutical trial. The changes that matter most in a person's life — the gradual integration of a new way of making meaning, the slow repair of a damaged relationship, the quiet emergence of a felt sense of belonging — unfold over years, not weeks. They are shaped by countless interacting factors that no randomized design can isolate. And they are often invisible to the instruments the outcome movement deployed.
The third wave: evidence-based practice and its shadow
By the early 2000s, the outcome movement had crystallized into a broader paradigm: evidence-based practice. The proposition was compelling. Fund and implement programs that have been shown, through rigorous research, to produce positive outcomes. Stop wasting resources on programs that have not been validated. In an era of scarce social funding and growing demands for accountability, evidence-based practice seemed like an unassailable commitment to both effectiveness and fiscal responsibility.
In many respects, it was. The evidence-based movement helped to expose programs that were ineffective or even harmful. It elevated the importance of evaluation in organizational culture. It created a shared language between funders, practitioners, and policymakers. These contributions are real and should be honored.
But the shadow of evidence-based practice is significant, and it is a shadow the Luminous approach takes seriously.
The hierarchy of evidence problem. Evidence-based practice established an implicit, and sometimes explicit, hierarchy of evidence, with randomized controlled trials at the top and qualitative, narrative, and experiential evidence near the bottom. This hierarchy privileged a particular kind of knowing — detached, objectified, quantitative — and systematically devalued other ways of understanding impact: the story a participant tells about how a program changed their sense of self, the felt shift in a community's collective energy, the somatic indicators a practitioner notices in a group.
The replication fallacy. The evidence-based paradigm assumed that if a program worked in one context, it would work in another, provided the program was implemented with fidelity. This assumption underestimates the profound influence of context. A parenting program that works beautifully in a middle-class suburb may fail utterly in a community shaped by generational poverty, institutional racism, and historical trauma. Not because the program is bad, but because the living system into which it is introduced is fundamentally different. Impact is not a property of a program alone. It is an emergent property of the relationship between a program and the system it enters.
The accountability trap. Evidence-based practice, in its most rigid forms, can create a culture of assessment as surveillance rather than assessment as learning. When organizations are evaluated primarily to determine whether they deserve continued funding, evaluation becomes a high-stakes performance rather than a genuine inquiry into what is happening and why. In such environments, organizations have powerful incentives to measure only what makes them look good, to select indicators they know they can hit, and to avoid asking the deeper questions that might reveal uncomfortable truths.
The mechanistic assumption. Perhaps most fundamentally, evidence-based practice inherited a mechanistic worldview from the medical model: the assumption that social interventions work like treatments, discrete inputs that produce predictable, linear effects. This assumption works reasonably well for simple problems. If children lack vaccinations, vaccinate them. But most of the challenges that matter most to human flourishing — poverty, loneliness, meaning-deficit, developmental stagnation, ecological destruction — are complex, adaptive, and irreducible to simple cause-and-effect chains.
The evidence-based movement, for all its contributions, did not ask the question the Luminous approach places at the center. What kind of consciousness is doing the measuring, and how does that consciousness shape what can be seen?
The fourth wave: developmental evaluation and complexity
In the early 2000s, a powerful counter-narrative began to emerge. Scholars and practitioners working in complex social systems — community development, systems change, organizational transformation — recognized that the existing evaluation paradigms were inadequate for the challenges they faced.
The most influential articulation of this counter-narrative came from Michael Quinn Patton, whose concept of developmental evaluation represented a paradigm shift. Patton argued that in complex, emergent situations, where the intervention itself is evolving, where the context is shifting, and where the outcomes cannot be predetermined, traditional evaluation is not merely insufficient but actively misleading. What is needed instead is an evaluative approach that:
- Serves learning rather than judgment. The primary purpose of developmental evaluation is not to render a verdict on a program's effectiveness, but to support ongoing adaptation and learning.
- Embraces emergence. Rather than measuring against predetermined outcomes, developmental evaluation tracks what actually emerges from the interaction between an intervention and its context, including outcomes that no one anticipated.
- Operates in real time. Rather than conducting evaluations after the fact, developmental evaluation embeds evaluative thinking within the ongoing process of innovation and adaptation.
- Honors complexity. Developmental evaluation recognizes that in complex adaptive systems, causation is non-linear, effects are distributed, and the relationship between action and outcome is often unpredictable.
Patton's work, along with contributions from scholars such as Patricia Rogers, Bob Williams, and Glenda Eoyang, opened the door to what might be called complexity-aware evaluation — approaches that take seriously the insights of complexity science, systems thinking, and developmental theory.
Around the same time, several complementary frameworks emerged. Outcome Harvesting, developed by Ricardo Wilson-Grau, offered a methodology for identifying and verifying outcomes in situations where they cannot be predicted in advance. Most Significant Change, developed by Rick Davies and Jess Dart, provided a participatory approach to evaluation that privileges the stories of those most affected by a program. Contribution Analysis, developed by John Mayne, offered a middle path between the impossibility of proving causation in complex settings and the abdication of any causal reasoning at all.
These approaches represent a genuine developmental advance. They honor the messiness of real-world change. They recognize that impact is not a fixed quantity to be measured but an ongoing, emergent phenomenon to be tracked, interpreted, and learned from. And they acknowledge what the earlier waves suppressed: that the observer is always part of what is observed, and that the act of measurement shapes the reality it claims to describe.
Yet even the complexity-aware wave has its limitations. Most of these approaches remain primarily cognitive. They analyze, interpret, and report, but they do not systematically engage the somatic, relational, and spiritual dimensions of impact. They expand the what of assessment — more kinds of outcomes, more kinds of evidence — without fundamentally questioning the who, the consciousness, the developmental stance, the embodied presence of the assessor.
What gets lost when only the countable is counted
It is worth pausing here to name, with as much precision as possible, what has been systematically excluded from the impact assessment story as it has unfolded.
Developmental depth. The vast majority of impact assessment frameworks measure change at the behavioral or attitudinal level. Did participants change their behavior? Did their attitudes shift? These are surface-level indicators of what may or may not represent a deeper transformation. A person who completes an anger management program and reports fewer angry outbursts has changed at the behavioral level. But has their relationship to anger changed? Have they developed the capacity to hold anger as an object of awareness rather than being subject to it? Have they moved from a meaning-making structure in which anger is an uncontrollable force to one in which anger is a signal to be attended to with curiosity? This deeper shift, the shift in how someone makes meaning, is invisible to conventional assessment tools, and it is precisely this shift that determines whether the behavioral change will endure.
Relational quality. Programs that serve communities, teams, and families inevitably affect the quality of relationships among participants. Yet relational quality — the degree of trust, mutuality, generative conflict capacity, and collective intelligence present in a relational field — is extraordinarily difficult to measure with conventional instruments. Surveys can capture perceptions of relational quality, but they cannot capture the living, felt reality of a relational field in motion. A team may report high satisfaction on a survey while harboring unspoken tensions that are eroding their collective capacity. Conversely, a team in the midst of a painful but transformative conflict may report low satisfaction while undergoing a relational deepening that will bear fruit for years.
Somatic knowing. The body carries information the mind cannot access. A practitioner who enters a community meeting and feels a constriction in the chest is receiving data about the emotional field of that community, data no survey can capture. Participants in a healing program may notice shifts in their somatic experience — the easing of chronic tension, the restoration of appetite, the return of dreamlife — long before they can articulate what has changed cognitively. The body is, in a very real sense, the first instrument of impact assessment. Yet it has been almost entirely absent from the field's methodology.
Ecological and systemic effects. Most impact assessment focuses on the immediate beneficiaries of a program. But programs exist within larger systems, and their effects ripple outward in ways that are often invisible to conventional frameworks. A leadership development program that transforms a single executive may, through that executive's changed behavior, affect the culture of an entire organization, the wellbeing of hundreds of employees, and the quality of service experienced by thousands of customers. These downstream effects are real, and they are almost never captured.
Spiritual depth and sacred participation. There is a dimension of human flourishing that transcends the psychological, the relational, and the systemic — a dimension that has to do with the sense of sacred participation in something larger than oneself. Call it meaning, call it purpose, call it the numinous. It is the quality that makes life feel not merely satisfactory but radiant. Programs that touch this dimension — contemplative practices, nature-based interventions, arts and creativity programs, grief rituals — produce effects that are real and profound and almost entirely unmeasurable by conventional means.
The assessor's own development. This is perhaps the most radical omission of all. The field of impact assessment has almost never turned its gaze upon the consciousness of the assessor. Yet the developmental stage, cultural assumptions, emotional state, and somatic awareness of the person doing the assessing profoundly shape what they can perceive. An evaluator operating from a conventional, achievement-oriented meaning-making structure will naturally privilege metrics that reflect achievement — efficiency, scale, cost-effectiveness. An evaluator operating from a more complex, integral meaning-making structure may perceive dimensions of impact the first evaluator literally cannot see. The instrument of assessment is not the survey or the interview protocol. It is the human being who designed, administered, and interpreted it.
The Luminous critique: why a new framework is needed
The Luminous approach to impact assessment does not reject the contributions of the waves that preceded it. Each wave — output counting, outcome measurement, evidence-based practice, developmental evaluation — represents a genuine advance, and each continues to offer tools and insights that have value. The Luminous approach practices transcend-and-include: honoring what each previous stage contributed while recognizing what it could not yet see.
But it insists that the field of impact assessment is ripe for a further developmental leap, demanded not by academic fashion but by the actual complexity of the challenges at hand. Climate change, mental health crises, the erosion of social cohesion, the search for meaning in a disenchanted world — these are challenges that cannot be adequately addressed, let alone assessed, by frameworks that reduce human flourishing to behavioral outcomes measured by standardized instruments.
What is needed is an approach that does six things.
Honors multiple forms of evidence. Quantitative data, qualitative narrative, somatic knowing, relational sensing, and contemplative insight are all legitimate forms of evidence. A progressive impact framework does not privilege one over the others but cultivates the capacity to integrate them, much as a skilled clinician integrates lab results, patient history, physical examination, and intuitive impression into a comprehensive diagnosis.
Includes developmental depth. Measuring what people do differently is important. Measuring how people make meaning differently is essential. A framework that cannot distinguish between surface-level behavioral compliance and deep structural transformation will consistently overestimate the impact of programs that produce the former and underestimate programs that cultivate the latter.
Assesses systemically. The unit of analysis for progressive impact assessment is not the individual alone. It is the nested system of individuals, relationships, organizations, communities, and ecosystems within which any intervention operates. Impact ripples through these nested systems in non-linear ways, and a progressive framework must be equipped to track those ripples.
Engages the body as an instrument of assessment. Somatic indicators — the felt sense of a group's energy, the physical markers of safety or threat in a room, the bodily signatures of transformation — are data. They are not sufficient by themselves, and they are essential. A framework that ignores the body's testimony is operating with a fraction of the available evidence.
Practices cultural humility. What counts as impact is not a neutral, universal category. It is deeply shaped by cultural values, worldviews, and power structures. A progressive impact framework asks whose definition of flourishing is being used, whose voices are centered in determining what success looks like, who benefits from the current metrics, and who is rendered invisible.
Turns the gaze upon the assessor. The Luminous approach takes seriously the recognition that the observer shapes what is observed. This means that progressive impact assessment includes practices for cultivating the assessor's own developmental capacity, somatic awareness, cultural humility, and ethical discernment. The most sophisticated assessment framework in the world, wielded by an assessor who lacks these capacities, will produce results that are technically precise and substantively hollow.
A developmental story, not just a historical one
Four waves, each with its contributions and its shadows. But this is not merely a story about the past. It is a story about the developmental trajectory of a field, one that mirrors in many ways the developmental trajectory of human consciousness itself.
The output era corresponds to what developmental theorists might call the concrete operational stage, the capacity to track tangible, visible, countable things. The outcome era corresponds to the formal operational stage, the capacity to think about causation, to construct logical models, to reason about things that are not directly observable. The evidence-based era represents the systematic stage, the commitment to rigorous methodology, standardized procedures, and replicable results. And the developmental evaluation era begins the movement into post-conventional consciousness — the recognition that reality is more complex than any single framework can capture, that context matters as much as content, and that the observer is always embedded in what is observed.
The Luminous approach invites a further step: the movement into what might be called integral assessment consciousness, a way of engaging with impact that holds multiple perspectives simultaneously, honors the body as well as the mind, includes the sacred as well as the secular, and recognizes that the deepest forms of human flourishing cannot be fully captured by any instrument yet devised.
This does not mean abandoning measurement. It means maturing in relationship to it — learning to hold metrics with the same reverence and humility with which a poet holds language, knowing that the words are never quite adequate to the experience, using them as skillfully as possible, always pointing beyond themselves toward the living reality they attempt to describe.
Common pitfalls and ethical cautions
Several cautions must be named before going further.
The romance of the unmeasurable. There is a temptation, particularly in contemplative and spiritual communities, to dismiss all measurement as reductive and to claim that the things that matter most are inherently beyond assessment. There is a kernel of truth in this — the deepest experiences of human life do resist full quantification — and it can also become a convenient excuse for avoiding accountability. The Luminous approach holds that most dimensions of human flourishing can be assessed, even if they cannot be reduced to a single number. The challenge is to develop assessment practices adequate to the complexity of what is being understood.
Measurement as control. Assessment always carries the potential to become an instrument of surveillance and control rather than learning and liberation. When assessment data is used primarily to reward or punish, to rank or sort, to justify or defund, it ceases to serve the purposes of human flourishing and becomes a tool of institutional power. The Luminous approach insists that assessment must always be in service to the people and communities being assessed, not merely to the funders or institutions that commissioned it.
The certainty trap. It is tempting to present assessment findings as definitive, as proof that a program works or does not work. But in complex human systems, certainty is almost always an overstatement. The Luminous approach practices epistemic humility: presenting findings as the best current understanding, acknowledging the limitations of the methods, and remaining open to being surprised.
Cultural imperialism in assessment. The dominant frameworks for impact assessment were developed primarily in Western, industrialized contexts. When exported to other cultural settings without adaptation, they can impose alien definitions of success, wellbeing, and flourishing. A progressive impact framework must always ask whether the framework is serving the people it claims to assess, or the cultural assumptions of those who designed it.
Mental health humility. Programs that address deep developmental, relational, or spiritual dimensions of human experience may surface difficult psychological material. Assessment practitioners must be aware of their own limitations and know when to refer participants to qualified mental health professionals. Assessment is not therapy, and the assessment process should never inadvertently cause harm.
Luminous invitations
Three things to notice as you read this chapter.
- What is your own relationship to measurement? Do you tend toward the pole of wanting everything quantified and proven, or toward the pole of resisting measurement as reductive? Neither pole is wrong, and noticing where you stand is the beginning of a more integrated relationship to assessment.
- What has been your experience of being assessed? Think of a time when you were evaluated, at school, at work, in a program. Did the assessment capture what was most important about your experience? What did it miss?
- Where in your body do you feel the tension between rigor and mystery? This is not a rhetorical question. The body carries real information about a person's relationship to knowing and not-knowing. Notice what arises.
Reflection questions
- Think about a program or intervention you have been involved in, as a designer, implementer, participant, or evaluator. What dimensions of impact were measured? What dimensions were ignored? What might have been different if the assessment had been more comprehensive?
- Consider the four waves of impact assessment described in this chapter. Which wave most closely reflects the assessment culture of your organization or community? What would it look like to take the next developmental step?
- This chapter argues that the consciousness of the assessor shapes what can be seen. How do you experience this in your own practice? What aspects of impact are you naturally attuned to? What aspects might you be blind to?
Practical exercise: your assessment autobiography
Take thirty minutes to write a brief assessment autobiography. Begin with your earliest memory of being measured or evaluated — a school test, a performance review, a medical examination. Trace the thread of assessment through your life. Notice the moments when assessment felt like a gift, when it helped you see something true about yourself. Notice the moments when assessment felt like a violation, when it reduced you to a number or a category that missed your essential humanity.
This exercise is not therapy. It is preparation. Progressive impact assessment begins with understanding one's own history with measurement — the biases, the wounds, and the gifts a practitioner brings to the work.
The next chapter moves from history to principles, articulating the five foundational commitments of the Luminous Impact Framework and what it means to build an assessment practice that honors the wholeness of human flourishing.
Output counting confuses activity with achievement.
The countable is seductive because it is easy to collect, easy to report and easy to defend.
When funding rides on the finding, evaluation stops being an inquiry and becomes a performance.
A framework blind to meaning-making will mistake compliance for transformation every time.
The body is the first instrument of impact assessment, and it has been almost entirely absent from the method.
The instrument is not the survey. It is the person who designed it.
Dismissing all measurement as reductive is a convenient way to avoid being accountable.
Certainty, in a complex human system, is almost always an overstatement.
Maturity in measurement is holding the metric the way a poet holds a word.
Works cited
- Patton, Michael Quinn. Developmental Evaluation: Applying Complexity Concepts to Enhance Innovation and Use. Guilford Press, 2010. The source of developmental evaluation as the chapter describes it.
- Scriven, Michael. “The Methodology of Evaluation.” Perspectives of Curriculum Evaluation, edited by Ralph W. Tyler, Robert M. Gagné and Michael Scriven, Rand McNally, 1967, 39–83. Source of the formative and summative distinction.
- Rogers, Patricia J. “Using Programme Theory to Evaluate Complicated and Complex Aspects of Interventions.” Evaluation, volume 14, number 1, 2008, 29–48.
- Williams, Bob, and Richard Hummelbrunner. Systems Concepts in Action: A Practitioner's Toolkit. Stanford University Press, 2010.
- Eoyang, Glenda H., and Royce J. Holladay. Adaptive Action: Leveraging Uncertainty in Your Organization. Stanford Business Books, 2013.
- Wilson-Grau, Ricardo. Outcome Harvesting: Principles, Steps, and Evaluation Applications. Information Age Publishing, 2018.
- Davies, Rick, and Jess Dart. The ‘Most Significant Change’ Technique: A Guide to Its Use. 2005.
- Mayne, John. “Contribution Analysis: An Approach to Exploring Cause and Effect.” ILAC Brief 16, Institutional Learning and Change Initiative, 2008.
- W. K. Kellogg Foundation. Logic Model Development Guide. 2004. The chapter describes logic models as emerging from the era without an originator; this is the most widely used statement of them, and is offered as a resolution rather than as the chapter's own citation.
- Piaget, Jean. The Psychology of Intelligence. Translated by Malcolm Piercy and D. E. Berlyne. Routledge and Kegan Paul, 1950. The chapter borrows the concrete operational and formal operational stages without naming him.
- Kegan, Robert. The Evolving Self: Problem and Process in Human Development. Harvard University Press, 1982. The subject-to-object shift the chapter uses when it asks whether a person can hold anger as an object of awareness rather than being subject to it, again without naming him.
- The randomized controlled trial as the gold standard of evidence-based practice. Described as borrowed from medical research; no single source is named, and none is needed.
- The parenting program that works in a middle-class suburb and fails in a community shaped by generational poverty, institutional racism and historical trauma. A constructed illustration rather than a case; no program, place or evaluation is named.
- The anger management participant who reports fewer outbursts. Likewise a constructed illustration, offered to distinguish behavioural change from a change in meaning-making.
- The Luminous Impact Framework, transcend-and-include, epistemic humility and integral assessment consciousness. House frameworks and terms of the Luminous Developmental Canon; internal to this book.