I am at a singular moment in my studies, focused on maturity models for AI governance (or the lack of it), a subject that keeps accumulating public cases of failed AI investments at large companies. The field is very new, and the market is still, in my view, crawling when it comes to maturity models. The FinOps Foundation itself acknowledges as much: it treats FinOps for AI as a discipline still taking shape, and its State of FinOps 2026 report records practitioners admitting that no one can yet answer whether AI is delivering the value it promised (FinOps Foundation, 2026). So I went back to academia, and I was caught off guard when I stumbled upon ‘Improving ratings’, Marilyn Strathern’s 1997 anthropological analysis of the British university audit system, including the 1996 Research Assessment Exercise.
The paper confronted me with statements economists know as Goodhart’s law:
“When a measure becomes a target, it ceases to be a good measure” (Strathern, 1997, p. 308).
She illustrates the point with British degree classifications: once a 2.1 (an upper-second-class degree) becomes the expectation, the grade stops discriminating between individual performances.
What is wrong with wanting a better score?
The reading is dense, anchored in an academic context I must admit I am not familiar with, but the paper’s closing comments made me reflect on what I learned during seven years at one of the Big Four: the consultant who genuinely goes deep into a company’s context, culture, and reality becomes far more capable of doing good work at the next client, quite unlike the consultant who spends that time figuring out how to transfer the model built at one client (the same model that fed the firm’s methodology and its maturity indicators) to all the others.
The classic example is the Toyota Production System, still current enough to warn us about the standardized application of maturity models in any field: decades of failed Lean Manufacturing implementations, because we consultants (and I include myself here) sold the visible tools, the kanban, the andon, the 5S, the kaizen, as universal “best practices,” without the tacit fabric that made them work at their origin.
Curiously, I lived this story from the other side of the counter, during a CMMI implementation at the software factory of a company that is now quite large: we reached the indicators and the levels, yet they did not translate into real maturity in the software we produced. They worked very well, however, as a “badge” for RFPs and commercial proposals. The score had become the target, and nobody there saw the problem, myself included.
What does this change for AI governance?
The paper’s somber conclusion made me rethink my strong intention of translating the FinOps Foundation’s Cloud maturity model into a scoring instrument for AI governance. Unfortunately, the temptation to turn a technical reference into a grading scale is enormous, precisely because grading scales sell well. The Foundation, evidently aware of the risk, states that an organization’s goal should never be simply to reach the Run stage in every capability, but to operate each one at the level its business context requires (FinOps Foundation, n.d.); and Rob Martin, in a provocatively titled piece on the Foundation’s blog, argues that maturity does not equal quality and that “there are no runners” (Martin, 2024). The model, therefore, should be used as a technical reference for diagnosis, exactly as the Foundation indicates, and not as maturity grades to be displayed.
Perhaps this is the warning that a 1997 paper about British universities was holding for those of us who work with maturity assessments in the corporate world: maturity models remain valuable instruments as long as they stay diagnostic references, and they tend to become corrupted the moment they turn into commercial requirements. If you run or commission this kind of assessment, I believe one uncomfortable question is worth asking before any engagement: are we measuring to decide better, or to display the score? My bet, and I admit it is still a bet, is that AI governance will only truly mature in companies that can answer that question without looking at their own marketing.
References
FinOps Foundation. (n.d.). FinOps maturity model. Retrieved July 29, 2026, from https://www.finops.org/framework/maturity-model/
FinOps Foundation. (2026). State of FinOps 2026 report. https://data.finops.org/
Martin, R. (2024, June 20). There are no runners. FinOps Foundation. https://www.finops.org/insights/no-runners/
Strathern, M. (1997). ‘Improving ratings’: Audit in the British university system. European Review, 5(3), 305–321. https://doi.org/10.1002/(SICI)1234-981X(199707)5:3<305::AID-EURO184>3.0.CO;2-4