Foreword
I have spent twenty years building teams that perform. The work has taken me through engineering, consulting, product, and the boardroom across three continents. The teams that lasted had two qualities in common: they were honest about what they did not know, and they had a way of working that did not depend on the person in charge.
The 4Ps came out of watching twenty-five teams adopt AI well and badly between 2023 and 2026. The good adoptions had four things present. The bad ones had at least one of those four missing. I started writing the pattern down because I was tired of explaining it from scratch each time.
I had already written a book about how teams compound performance. That book, Team Performance Uplift, lives around six dimensions, the 6Ps. It explains how groups of people lift each other. It does not address what AI does to the contract between those people, the work, and the customer. AI changed that contract in three ways. First, it collapsed the cost of work, in a three-step arc: drafts, then tools, then autonomous action, each at lower cost per cycle and higher autonomy than the last. Second, it shifted accountability from the author to the operator. Third, it introduced a new class of risk that is invisible until it lands in front of a regulator or a customer. Part X picks up the autonomous-action end of that arc explicitly.
The 4Ps is the answer to those three shifts. It is operator-first. It is written for the leader who has to make AI productive on Monday morning, profitable in the quarter, and explainable on Friday afternoon. It is not a compliance manual, although it will satisfy compliance. It is not a moral treatise, although it has a moral spine. It is a working framework for responsible growth.
Together, the 4Ps and the 6Ps form an operating system. The 6Ps explains how teams compound. The 4Ps explains how AI in those teams stays trustworthy. If a team is using the 6Ps and ignoring the 4Ps, they are scaling something that has not been governed. If a team is using the 4Ps and ignoring the 6Ps, they have governance and no engine. The two were always meant to sit alongside each other.
A note on the title. I considered calling this Responsible AI for Operators and decided against it. Responsible AI as a phrase has been claimed by enough vendors that it has lost its working edge. I considered calling it AI Governance That Works. That is closer, but it asks the reader to take the title's word for it. The 4Ps is a description. It names what is in the box. Primed, Principled, Practised, Protected. A leader can read those four words and tell whether the framework is offering something they already have or something they do not.
A note on the audience. The framework is written for the operator. By operator I mean the person who is accountable for an outcome, has a budget, has a team, and has to make decisions on incomplete information. Operators read differently from regulators and academics. They read for what to do next. The framework is structured accordingly. Each chapter ends with a diagnostic checklist. Each P has a maturity rubric. The 30-day rollout in Part VI is a working plan, not an illustrative example.
A note on tone. I have tried to keep the writing honest about uncertainty. AI moves quickly. The regulatory landscape moves slowly but unpredictably. The 4Ps is built to absorb both. Where a sub-dimension might change as the technology shifts, the framework names the change point rather than pretending it does not exist.
Use the framework. Adapt it. Argue with it where you need to. If it helps you ship better AI more responsibly, it has done its job.
Segun Osu London, May 2026
Part I: Why a new framework was needed
Most AI governance frameworks are written for the regulator, not the operator, and not with the creation of business value in mind. They start with risk classification, move to control objectives, and arrive at a policy document the team rarely opens. They tell a business what to comply with. They do not tell it how to do AI in a team on Monday morning. They do not highlight how good governance creates business value.

The 4Ps starts from the other end. It starts with the operator. Specifically, with the leader who is asked three questions at the same time: how do we get more value out of AI, how do we make sure it is safe, and how do we prove that to the people who are asking. The 4Ps gives that leader a single answer.
There are good frameworks already in the world. The NIST AI Risk Management Framework is rigorous and well-structured. The EU AI Act has given Europe a legal floor. ISO 42001 sets a standard for an AI management system. OECD principles give a values orientation. The Oxford curriculum on AI governance covers all of these. Each is useful. None is operator-first.
Operator-first means three things. First, the framework can be deployed by a team without a dedicated risk function. Second, the framework produces decisions, not documents. Third, the framework integrates with how teams already work, rather than replacing it.
The need for an operator-first framework grew sharper between 2023 and 2026. The first wave of enterprise AI adoption taught the market a series of lessons that the regulator-first frameworks did not predict. Pilots stalled because no-one had decision rights. Models were deployed without anyone holding the operational risk. Vendors changed terms mid-contract and the procurement teams did not notice. Employees used shadow tools because the official ones were too slow. Each of these is an operator problem. Each is what the 4Ps is built to prevent.
The 4Ps does not replace NIST, the EU AI Act, ISO 42001, or the OECD principles. It sits on top of them. It uses them where they apply. It points at them in the Protected dimension. The contribution of the 4Ps is not new regulation. It is a way of doing AI inside a team that satisfies regulation without being run by it.
The framework name reflects its sequence. Primed before Principled because there is no point writing principles for a team that cannot use the tools. Principled before Practised because practising without principles produces drift. Practised before Protected because protection without practice produces theatre. Each P assumes the one before it. Each P is testable. Each P maps to specific work a leader can commission this quarter.
There is a second reason a new framework was needed. The gap between AI strategy decks and AI operating reality has widened in three years. The decks describe what is possible. The operating reality is shaped by data that is too dirty for the use case, models that, even with retrieval-augmented generation and grounding controls in place, can still hallucinate at the moments customers notice most, vendors who change their terms when the spend grows, and employees who route around official tools because the official tools are too slow. None of these problems are solved by a regulator-first framework. All of them are solved by operating discipline.
A third reason. The teams that adopted AI well between 2023 and 2026 were not the ones with the largest budgets. They were the ones with the most disciplined sequencing. They built literacy before they bought platforms. They wrote principles before they shipped pilots. They put a lifecycle around their first use case before they added their second. They rehearsed incident response before they had an incident. The pattern was visible across companies of very different sizes, in very different industries, in very different regulatory environments. The 4Ps is the codification of that pattern.
A fourth reason. Most AI failures are not model failures. They are process failures, decision-right failures, or vendor failures, dressed up as model failures in the post-mortem. A hallucinated answer in a customer email becomes an incident not because the model hallucinated, residual hallucination exists even with retrieval-augmented generation and structured-output controls, but because no-one decided who reviews customer-facing AI output, and no-one specified the human-in-the-loop rule for that channel. The model layer's controls reduce the rate. The operating layer catches what slips through. A framework that addresses only the model layer will not catch this kind of failure. The 4Ps addresses the layers around the model deliberately.
The argument for an operator-first framework is therefore practical, not ideological. The regulator-first frameworks are necessary. They are not sufficient. The operator needs a working frame that produces decisions, creates business value, fits into the team's existing rhythm, holds up in front of an auditor, and improves over time. That is the brief the 4Ps was written against.
Part II: The genesis of the 4Ps
The 4Ps was shaped by three continents working as equal contributors, two universities, one philosophy training, and roughly twenty-five client engagements. None of that would matter if the framework did not work. The reason for naming the sources is that each one contributed a specific element, and a reader who wants to deepen any dimension can follow the trail back to where it came from. The three continental lenses below are presented in the order they entered the framework's evolution, not in any order of weight. Each is load-bearing.
The European chapter contributed the rhythm of operating inside mature, regulated institutions. Sixteen years split between the United Kingdom and Italy meant working in settings where regulators had real teeth, audit cycles were long, and the gap between a process on paper and the same process on a Monday morning was visible every week. Italian commercial practice is unusually attentive to who actually does the work, as distinct from who is responsible for the work in theory. The Italian language itself distinguishes the responsabile (the named owner) from those who fanno il lavoro (the people who actually move the work forward), and well-run Italian organisations watch both. That habit fed the Practised dimension. Practice is what gets done when no-one is in the room.
The African chapter contributed the discipline of relational accountability. In Yoruba practice, decisions are weighed against the elders who will hear about them and the descendants who will live with them. A decision is not a private act. It is a public commitment carried by a name. That instinct shows up in the Principled dimension. Principles are not slogans. They are commitments a named person stands behind, and the name follows the decision long after the moment has passed.
The transatlantic commercial chapter sharpened the value orientation. Selling research to financial services in New York and selling product strategy to multinational corporates in London is the same skill performed in different accents. Both require translating a complex idea into a number a buyer can defend internally. That habit feeds across every P. Good governance is governance that pays back.
The three lenses cross-correct each other. The Italian habit of watching real practice keeps the Yoruba habit of public commitment from becoming theatrical. The transatlantic habit of defensible numbers keeps both from drifting into culture as a substitute for results. Stripped of any one of the three, the framework would be poorer.
The two universities sit on either side of the 4Ps. Oxford, through the AI Governance programme at Saïd Business School in December 2025, provided the regulatory landscape, risk classification, governance structures, and ethics frameworks. That work feeds P2 (Principled) and P4 (Protected). The 100% grade is less interesting than the curriculum itself, which forced a structured engagement with the EU AI Act, the NIST AI Risk Management Framework, and the OECD principles in a single coherent body of work.
Wharton, through the AI for Business Specialization completed in March 2025, provided the commercial frame. Four certificates - AI Strategy and Governance (100% grade), AI Fundamentals for Non-Data Scientists, AI Applications in Marketing and Finance, and AI Applications in People Management - covered the value side. Where Oxford asked how to govern AI, Wharton asked what to do with it. That work feeds P1 (Primed) and P3 (Practised).
The Practical Philosophy training is the quiet contributor. Years of weekly study at the School of Practical Philosophy - Plato, the Vedic traditions, critical thinking habits - taught a working discipline. Decisions compound. Rules outlive personalities. The same question, asked of a different mind, returns a different answer. The Principled dimension borrows directly from this training. It is the reason the framework treats written principles as load-bearing, not decorative.
The engagements supplied the evidence. A 45% performance lift and a 15% cost reduction at Haleon PLC came from disciplined practice plus a refusal to deploy AI in places where the operating data was not yet ready. A 40% lift at GSK PLC came from the same template applied earlier. Ogier Group, working in a regulated legal and financial services environment, taught what corporate and data governance look like when failure has a regulator at the end of it. Teamsmiths AI SaaS, the proprietary uplift product, demonstrated a 45%+ uplift in team throughput and roughly 99% delivery predictability in three months, using GenAI large language models, retrieval-augmented generation, and the proprietary uplift model. Five production agentic AI platforms - Revenue-Risk-Radar, Songita, Deputee, SDRAgent, AIMarketer, and Piccolingo - taught what shipping looks like, as distinct from what advising looks like.

Earlier work supplied calibration. A £25M accounts receivable situation at Philips Italy provided the operational evidence behind a project management application at Philips that subsequently safeguarded £22M annually. Managing technology planning and critical EU data migrations at Reuters. Designing strategy for Orange, VW, and ICI at IF Consulting in support of $3.2bn new revenue. Launching a financial analytics product at Knight Ridder/McClatchy that generated £1m in year one. Designing the strategic planning process and governance framework that enabled an Executive Agency transition for the UK Government. Delivering sections of £20M to £25M infrastructure projects at Balfour Beatty. Growing a portfolio from $3M to $10M over five years and halving delivery times as Senior Product Director, EMEA Digital Media and eCommerce at Gartner, recognised with the Chairman's award. Co-founding G2Guide from 2005 to 2020 with partnerships at Reed Elsevier and IDOX, delivering up to 53% cost savings on supplier contracts.
The pattern in all of it is the same. Governance that ignores the operator fails on contact. Practice that ignores governance scales the wrong thing. The 4Ps is the version of that lesson that fits on a page.
A word on why the framework has four dimensions rather than five, six, or three. Early drafts tested versions with three Ps (Primed, Practised, Protected) and five Ps (adding a Performance dimension between Practised and Protected). The three-P version was operationally tight but lacked the moral spine that the Principled dimension provides. The five-P version repeated material from the 6Ps in Team Performance Uplift and confused readers about which framework was in use. Four Ps emerged as the version that held under operating stress without overlapping the companion framework.
A word on what the 4Ps deliberately does not contain. It does not contain a model evaluation methodology. There are excellent published methodologies for evaluating model performance, fairness, and robustness, and they evolve faster than any governance framework should. The 4Ps points at evaluation methodology in P4 (Protected) without prescribing one. A team should pick the evaluation methodology that fits its risk class and its industry.
It does not contain a procurement template. AI vendor contracts are too varied and too jurisdiction-specific to be templated usefully. The 4Ps specifies what a vendor relationship must produce (in P3 Practised, the vendor and partner discipline sub-dimension) without specifying the contractual form.
It does not contain a training curriculum. Workforce literacy is a sub-dimension of P1 (Primed), but the actual training material is role-shaped and industry-shaped. The framework specifies the outcome required without dictating the syllabus.
The omissions are deliberate. A framework that tries to be both an architecture and the implementation of that architecture is too heavy to deploy. The 4Ps is the architecture. The implementation is shaped by the team using it.
Part III: The framework in one page
The harness, not the cage. Most AI governance is sold as a cage. A cage stops the animal. It is built to prevent and to fence, and a caged capability does no work. Worse, a team that cannot use AI safely will in the end use it unsafely, around whatever fence you build.
The 4Ps is a harness. A harness does three things a cage never will. It transfers raw power into useful work. It hands you the reins, so you choose the direction. It rigs a safety line, so one bolt does not end you. AI is power. Left loose it bolts, breaks things, or stands idle because nobody trusts it near the business. Strap on the 4Ps and that same power pulls a load, in a direction you set, with a line that catches you when it spooks.
The four parts map onto the parts of a real harness, which is what makes the framework easy to teach and hard to forget.
Primed is the fitting. You cannot harness an animal you have not sized up. Skip it and the harness chafes, the rules do not fit the work, and the team throws the whole thing off.
Principled is the reins. Direction and control. Who decides where this pulls, and who is allowed to stop it.
Practised is the pull. Power transferred to the load, day after day. The 6Ps operating model is the motor that lives inside this part.
Protected is the safety line. The strap that catches you when something bolts. It also covers the animal that stalls rather than runs: a vendor rationing compute is a failure you rig for in advance, because AI capability is a supply-constrained input, not an elastic utility.
Hold that image and the rest of the book reads as one idea. Govern AI like a harness, not a cage.

The 4Ps are four dimensions a team must hold together to do AI well. They are sequential when first introduced and continuous when in operation. A diagram of the framework appears in Appendix C as the 4Ps wheel.
P1 - Primed. The organisation has the knowledge, mindset, data, and infrastructure to use AI well. AI literacy and readiness across leadership, workforce, data estate, and platforms. Sub-dimensions: Leadership literacy. Workforce literacy. Data readiness. Infrastructure readiness. Use-case clarity.
P2 - Principled. Ethics, accountability, and governance written down and lived. Decision rights are explicit. Ethics review has a process. Transparency standards apply to internal and external use. Human-in-the-loop rules are documented. Sub-dimensions: Stated principles. Decision rights (clear RACI). Ethics review process. Transparency standards. Human-in-the-loop rules.
P3 - Practised. AI is operationalised through disciplined practices day to day. Use-cases follow a lifecycle. Models follow a lifecycle. Teams operate to a cadence. Capability is built deliberately. Vendors are held to a standard. The 6Ps from the Team Performance Uplift book is the operating model that lives inside P3, embedded as the Team capability sub-dimension. Sub-dimensions: Use-case lifecycle. Model lifecycle. Operating cadence. Team capability (6Ps integration). Vendor/partner discipline.
P4 - Protected. Risk management, compliance, audit, and security. The risk inventory is current. Regulatory alignment with the EU AI Act, NIST AI RMF, and ISO 42001 is documented. Model cards are produced for material models. Audit trails are usable. Incident response is rehearsed. Sub-dimensions: Risk inventory. Regulatory alignment (EU AI Act, NIST AI RMF, ISO 42001). Model cards. Audit trail. Incident response.
The order matters. Primed before Principled because principles drafted by a team that cannot use the tools will be unrealistic. Principled before Practised because practice without principle drifts. Practised before Protected because protection layered on top of weak practice produces audit theatre.

The four together form a wheel that turns continuously, not a checklist that finishes.

Reading the one-page summary.
The one-page summary is designed to be the working artefact a leader keeps to hand. Three uses dominate in practice.
The first use is induction. A new starter on the team reads the one page in five minutes and has a working frame for everything the team does with AI. The artefacts they will encounter - principles, RACI, lifecycles, model cards, risk inventory, audit trail - all live inside one of the four Ps.
The second use is conversation. When a stakeholder asks a hard question - is this safe, is this compliant, is this measured - the one page is the answer's structure. The framework names the dimension in which the answer lives. The detail comes from the artefact inside that dimension.
The third use is prioritisation. When the team has a quarter's worth of capacity and ten things they could do, the one page is the lens through which the ten are sorted. Items that sit in the lowest-scoring dimension are prioritised. Items that sit in the highest-scoring dimension are deferred.
A team that adopts the framework and never looks at the one page again has missed the point. A team that prints the one page, pins it on the wall, and uses it weekly has understood it.
Part IV: Each P in detail

Chapter 1: P1 - Primed
What P1 covers. Primed is the dimension that asks whether the organisation has the knowledge, mindset, data, and infrastructure to use AI well. It is a prerequisite for everything that follows. A team that is not Primed will produce principles it cannot honour, practices it cannot sustain, and protections it cannot operate. Primed is the soil. The other three Ps are what grows in it.
Sub-dimensions.
Leadership literacy. The executive team can describe in plain language what generative AI is, what it can and cannot do, what it costs, and where it fails. They can read a model card. They can sit through a vendor demo without being either dazzled or dismissive. They can ask the second question.
Workforce literacy. Employees at every level have a working understanding of AI relevant to their role. A customer service agent knows when to trust a generated reply. A finance analyst knows when a model is hallucinating a number. A product manager knows the difference between a foundation model and a fine-tune. Literacy is role-shaped, not generic.
Data readiness. The data the organisation needs in order to use AI is findable, governed, accurate enough for the use case, and legally available for the use case. Data readiness is not a perfect-data target. It is a fit-for-use target.
Infrastructure readiness. The platforms, pipelines, and tooling the organisation needs are in place. This includes access to models, vector stores, observability tools, and the security posture around them. Infrastructure readiness includes the unglamorous parts: identity and access management, logging, cost monitoring. It also includes supply headroom and a fallback the organisation controls, so that a vendor running short of compute does not become an outage the team cannot answer.
Use-case clarity. The organisation has a short list of AI use cases it cares about, each described in terms of business outcome, owner, success measure, and risk class. A long list of vague aspirations is a sign of low Primed maturity. Purpose at the use-case level lives here. Organisation-level Purpose, the why behind doing AI at all, sits upstream of the framework in the team's strategy. Team-level Purpose lives in the 6Ps inside Practised.
What good looks like. A team rated high on Primed can produce, within a working day, a one-page summary of its top three AI use cases, the data each one depends on, the platforms each one uses, and the named person accountable for each. Leadership can answer questions about each in plain language. The workforce has been trained for the use cases relevant to their roles. Data and infrastructure decisions are aligned with the use cases, not bought ahead of them.
Typical failure modes. Buying a platform before deciding the use case. Training the whole organisation in generic AI literacy and skipping role-shaped depth. Treating data readiness as a separate three-year programme rather than a per-use-case requirement. Confusing executive enthusiasm with executive literacy. Allowing shadow tools because the official platform is not fit for purpose.
Diagnostic checklist for P1.
Can each member of the executive team describe the top three AI use > cases in plain language?
Has the workforce been trained in AI relevant to their roles within > the last twelve months?
Is the data each priority use case depends on documented and > fit-for-use?
Is the platform stack the organisation uses for AI documented, with > cost and access controls in place?
Is there a written list of priority AI use cases, each with an owner > and a success measure?
Are there approved channels for employees to raise new AI ideas, and > are those channels active?
Is shadow AI use measured, and if so, is it trending down as > official tools improve?
Source. P1 draws principally from the Wharton AI for Business Specialization, in particular AI Fundamentals for Non-Data Scientists and AI Applications in People Management. Leadership and workforce literacy as deliberate disciplines are Wharton-shaped. Use-case clarity and data readiness are operator-shaped, drawn from the Haleon and GSK engagements where the largest cause of stalled AI work was an unclear use case rather than an unfit model.
Field note. A consumer goods business asked for help in 2024 with what they described as a stalled AI programme. They had bought three platforms and trained six hundred employees in a generic literacy course. Six months in, none of the pilots had moved into production. The diagnostic took half a day. The literacy course had been generic, not role-shaped, so the marketing team could not tell whether the model was the right tool for the job, and the data team could not tell whether the data was fit for the use case. The platforms had been chosen before the use cases were named, so the use cases that mattered most needed a fourth platform. The pilots were stalled because no-one owned them. We rebuilt the literacy curriculum in two role-specific tracks. We named owners for the three priority use cases. We retired one platform. The first use case moved into production inside ten weeks. The lesson is in the order. The team had skipped from problem statement to platform purchase without passing through the Primed checklist.
Field note. A regulated services business in 2025 reached an unusual state: high data readiness, low leadership literacy. The data team had spent two years building a clean estate. The board could not explain in plain language what generative AI was. The result was a series of pilots that the board approved without understanding, and a data team frustrated that their best work was being deployed without governance. The intervention was a one-day executive workshop, run twice, with role-shaped material rather than vendor demonstrations. Within a quarter, board approval of AI use cases shifted from generic green-lights to specific decisions with named risk class. The data team's work was finally being used on the problems it had been prepared for.
Chapter 2: P2 - Principled
What P2 covers. Principled is the dimension that asks whether ethics, accountability, and governance are written down and lived. It is the dimension that survives a change of leadership. Principles that exist only in a leader's head die with the leader's tenure. Principles that are written, agreed, and operated outlive the people who drafted them.
Sub-dimensions.
Stated principles. The organisation has a short, public set of principles for how it uses AI. The set is short because long lists are not lived. The set is public because public commitment changes behaviour. Typical principles cover fairness, transparency, accountability, privacy, safety, and human oversight, adapted to the organisation's context.
Decision rights. For each material AI decision, the organisation knows who is Responsible, who is Accountable, who is Consulted, and who is Informed. The RACI is published. RACI without publication is a draft.
Ethics review process. There is a defined process for raising and resolving ethics concerns about an AI use case before it goes live, and a process for revisiting concerns after it goes live. The process has a chair, a quorum, and a meeting cadence. It does not need a large committee. It needs a working one.
Transparency standards. The organisation has clear rules about when AI use must be disclosed to a customer, an employee, a partner, or a regulator. The rules apply to internal and external use, not just customer-facing use.
Human-in-the-loop rules. For each AI use case, the organisation has decided whether a human reviews each output, samples outputs, reviews outputs in defined edge cases, or does not review at all. The rule is matched to the risk class of the use case.
What good looks like. A team rated high on Principled can show the principles document, the RACI for each material AI decision, the ethics review minutes from the last quarter, the transparency rules with examples of where they have been applied, and the human-in-the-loop rule for each priority use case. New starters meet the principles in their induction. Vendors are required to acknowledge them.
Typical failure modes. A principles document that no-one references. A RACI that confuses Accountable with Responsible, so two people own a decision and neither owns it. An ethics review process that meets twice and then dissolves. Transparency rules that apply to customers but not employees, or to external partners but not internal stakeholders. Human-in-the-loop rules that are uniformly cautious or uniformly absent, regardless of risk class.
Diagnostic checklist for P2.
Are there written AI principles, agreed by leadership, signed by a > named accountable executive?
Is there a published RACI for material AI decisions, including model > selection, deployment, and retirement?
Has the ethics review process met at least once in the last quarter, > with minutes?
Are transparency standards documented for internal and external AI > use?
Are human-in-the-loop rules matched to the risk class of each use > case?
Are vendors required to acknowledge the principles as part of > contracting?
Are new starters introduced to the AI principles in their induction?
Source. P2 draws principally from the Oxford AI Governance programme at Saïd Business School. The structural elements - decision rights, ethics review, transparency, human oversight - map directly onto the OECD principles, the EU AI Act's obligations on high-risk systems, and the governance structures explored at Oxford. The Practical Philosophy training contributed the disposition that principles are load-bearing rather than decorative. The Ogier Group engagement contributed the field experience of operating principled governance in a heavily regulated environment, where the regulator does not accept the answer that the principles were never tested.
Field note. At Ogier Group, the most useful intervention was unglamorous. The principles existed. The RACI existed. They did not connect. A model selection decision sat with the data team. A deployment decision sat with the product team. A risk classification decision sat with compliance. None of them met as a group. The result was a series of decisions made in isolation, with the inevitable consequence that an upstream decision foreclosed a downstream one. We instituted a fortnightly thirty-minute decision forum with named seats, a published agenda, and a written decision log. Within two months, the same three teams were producing decisions in half the time, with full audit trail, and the regulator's first-round questions were answered in a single email.
Field note on principles. A common mistake is writing too many principles. Twelve principles is a brochure. Five principles is a working frame. The five should usually cover fairness, transparency, accountability, privacy, and human oversight. Add safety if the use case is safety-critical. Add explainability if the use case is in financial services or healthcare. Resist adding more. Each additional principle dilutes the others and makes the set harder to operate.
Field note on RACI. The most useful RACI is one that distinguishes Accountable from Responsible cleanly. Accountable is the single named person who carries the decision if it goes wrong. Responsible is the named person or team who does the work. A RACI that has two people listed as Accountable for the same decision will produce inaction when the decision is hard. A RACI that does not name a person at all, only a role, will produce inaction when the role-holder changes.
Chapter 3: P3 - Practised
What P3 covers. Practised is the dimension that asks whether AI is operationalised through disciplined practices day to day. Principles without practice are wishful thinking. Practice without principles is undirected motion. P3 is where the framework meets the team.
This is the dimension that integrates the 6Ps from Team Performance Uplift. The 6Ps - People, Purpose, Principles, Practices, Productivity, and Performance - is the operating model that sits inside P3 as the Team capability sub-dimension. A team that is doing the 6Ps well is producing the conditions inside which the 4Ps' Practised dimension can succeed.
Sub-dimensions.
Use-case lifecycle. Each AI use case moves through a defined lifecycle: idea, scoping, build, pilot, scale, monitor, retire. Each stage has entry and exit criteria. The lifecycle is short enough to be remembered and detailed enough to be operated.
Model lifecycle. Each model the organisation uses, whether internal or vendor-supplied, has a defined lifecycle: selection, sandbox evaluation, pre-deployment red teaming, deployment, continuous monitoring for performance and drift, retraining or refresh, retirement. The model lifecycle integrates with the use-case lifecycle. Red teaming is not optional for high-risk classes. Continuous monitoring is not optional for any production model.
Operating cadence. The organisation runs AI work to a deliberate rhythm. Weekly operational reviews. Monthly performance reviews. Quarterly portfolio reviews. The cadence is published, attendance is named, and decisions are recorded.
Team capability (6Ps integration). The team has the People, Purpose, Principles, Practices, Productivity, and Performance in place to sustain the AI work. The 6Ps is the operating model. The 4Ps is the governance frame.
Vendor and partner discipline. The organisation has a defined approach to selecting, contracting, and managing AI vendors. The approach covers technical due diligence, data handling, indemnity, performance commitments, exit, and the alignment of vendor practices with the organisation's principles. Performance commitments include deliverability under scarcity: whether the contract buys committed throughput or capacity reservation rather than best-effort terms, what priority tier applies when the vendor rations, and whether the vendor is diversified across silicon, cloud, and power. A vendor that cannot guarantee supply is a continuity risk, not only a commercial one. Vendor discipline is not a single procurement decision. It is an ongoing relationship.
What good looks like. A team rated high on Practised can show the use-case lifecycle stages with a current example at each stage. It can produce model cards for each material model in production. It runs an operating cadence with named attendance and recorded decisions. It can demonstrate that the 6Ps practices are present in the team's daily work. It has a vendor scorecard that is updated quarterly.
Typical failure modes. A use-case lifecycle that exists in a diagram but not in the team's calendar. A model lifecycle that covers selection well and retirement badly, so retired models continue to run silently. An operating cadence that turns into status theatre rather than decision-making. Team capability that is described as a training problem rather than an operating model problem. Vendor relationships that are managed through procurement at contracting and forgotten thereafter.
Diagnostic checklist for P3.
Is there a documented use-case lifecycle with entry and exit > criteria for each stage?
Is there a documented model lifecycle, with named owners at each > stage?
Does the operating cadence have a published rhythm, named > attendance, and recorded decisions?
Are the 6Ps - People, Purpose, Principles, Practices, Productivity, > Performance - in place as the operating model?
Are vendor relationships managed against a written scorecard updated > at least quarterly?
Are retired models actually retired, including the removal of access > and the archiving of audit trails?
Are practitioners given dedicated time to learn and apply new AI > techniques, separate from delivery time?
Source. P3 draws from the Wharton AI Applications courses (Marketing, Finance, People Management) for the use-case taxonomy, the Haleon and GSK engagements for the lifecycle discipline and the operating cadence, and the Team Performance Uplift 6Ps work for the team capability operating model. The five production agentic AI platforms (Revenue-Risk-Radar, Songita, Deputee, SDRAgent, AIMarketer, Piccolingo) contributed the vendor discipline learning, in particular what happens when a model provider changes its terms mid-project.
Field note. At Haleon PLC, the practice that produced the largest single improvement was a weekly operational review with named attendance, a published agenda, and a written decision log. Before the change, AI use-case decisions had been made in ad-hoc forums, with the consequence that decisions made on Tuesday were unmade on Thursday by people who had not been in the Tuesday room. After the change, the same number of decisions were made, but each one held. Performance lifted because effort stopped being wasted on rework. The 45% performance lift across product, engineering, and data teams is partly a model story, but it is mostly a cadence story.
Field note on the use-case lifecycle. The stages do not need to be sophisticated. Idea, scoping, build, pilot, scale, monitor, retire. What matters is that each stage has entry and exit criteria, and that the criteria are honoured. An idea moves to scoping only when a sponsor has been named. A pilot moves to scale only when the success measure has been hit on real users. A use case moves to retire when its purpose is served or its risk has changed. The most common failure is no retirement, which produces a sprawl of half-active use cases that no-one is monitoring.
Field note on vendor discipline. A model provider changing its terms mid-project is the single most common vendor problem in AI. The mitigation is not a clever contract clause. It is a deliberate dependency map showing which use cases depend on which vendor and what the switching cost is. The dependency map is reviewed each quarter. When a vendor changes its terms, the response is shaped by the dependency map, not by panic.
Chapter 4: P4 - Protected
What P4 covers. Protected is the dimension that asks whether risk management, compliance, audit, and security are operating. It is the dimension that proves the previous three are working. A team that cannot answer a regulator, a customer, or an internal auditor with current evidence is a team that has not yet earned the Protected rating, regardless of how mature the other Ps are.
Sub-dimensions.
Risk inventory. The organisation has a list of AI risks specific to its use cases, with each risk classified by impact and likelihood, with owners and current mitigations. The inventory covers technical risk, ethical risk, regulatory risk, security risk, operational risk, and supply risk, the last being the vendor's ability to deliver the compute and throughput a use case depends on and the concentration that builds when critical work runs on a single supplier. High-risk use cases have a documented Algorithmic Impact Assessment (an AIA, the AI cousin of a DPIA) sitting alongside the risk entry. The inventory is reviewed at the operating cadence, not maintained as a static document.
Regulatory alignment. The organisation's AI work is aligned with the regulations that apply to it. For most readers, that means the EU AI Act, the NIST AI Risk Management Framework, and ISO 42001 at minimum, with sector-specific regulation layered on top. Alignment is documented in a mapping that any new starter can read.
Model cards. Each material model in production has a model card describing its purpose, training data, performance, limitations, intended use, and the named owners. Model cards are public to internal stakeholders, accessible to auditors, and updated when the model changes.
Audit trail. The organisation can reconstruct, for any AI-mediated decision of material consequence, what the model did, what the human did, when, and why. The audit trail is usable in practice, not just present in theory.
Incident response. The organisation has a defined process for what to do when an AI system fails, produces a harmful output, breaches a control, attracts regulatory attention, or is rationed by a supply-constrained vendor. The process has been rehearsed, including the day a vendor throttles or withdraws capacity and work moves to a fallback the organisation controls. The escalation chain is current.
What good looks like. A team rated high on Protected can produce, on request, the AI risk inventory with current mitigations, the regulatory mapping with current obligations, the model cards for production models, an audit trail extract for a recent decision, and the incident response playbook with evidence of at least one rehearsal in the last twelve months. The Protected dimension is what allows the organisation to grow AI use without growing AI risk faster than its ability to manage it.
Typical failure modes. A risk inventory that was produced for an audit and not maintained afterwards. Regulatory alignment expressed as a wall of text without a mapping back to specific controls. Model cards that exist but are not updated when the model changes. Audit trails that are technically present but not retrievable inside an incident window. Incident response that has never been rehearsed, so when an incident happens the playbook is opened for the first time.
Diagnostic checklist for P4.
Is the AI risk inventory current, with each risk owned by a named > person?
Is there a written mapping from the organisation's AI work to the > EU AI Act, NIST AI RMF, and ISO 42001?
Does every material production model have a current model card?
Can the organisation reconstruct an AI-mediated decision of material > consequence from its audit trail?
Has the incident response playbook been rehearsed in the last twelve > months?
Are there named accountabilities for AI risk at the executive level?
Does the audit trail include human override events as well as model > decisions?
Source. P4 draws principally from the Oxford AI Governance programme - the regulatory landscape, risk classification, and audit posture - and from the Ogier Group engagement, where corporate and data governance had to operate against a real regulator in a real industry. The Practical Philosophy training contributed the disposition that the test of governance is what happens when it is challenged, not what it says on a slide.
Field note on model cards. Model cards are not optional. The convention has matured to the point that any production model deployed without a model card is below the standard the regulators and the larger customers expect. The model card does not need to be long. Two pages is usually enough. The card states the purpose of the model, the intended use cases, the data the model was trained on, the model's known limitations, the human-in-the-loop rule, and the named owners. The card is versioned, and the version history is preserved.
Field note on incident response. The first AI incident a team encounters is rarely the kind anticipated by the playbook. A model produces an inappropriate output in a customer channel. A vendor announces a change of terms with a thirty-day notice. A regulator publishes new guidance that retroactively re-classifies a deployed use case. A data leak originates from a fine-tuning dataset. Each of these has happened. The playbook does not need to enumerate every scenario. It needs to define the first hour of response: who is called, what is documented, what is paused, and who decides what to communicate externally.
Field note on audit trail. Audit trails fail in practice for two reasons. They are not retrievable inside an incident window, or they capture model decisions but not human override events. A retrievable audit trail can produce, within an hour, the decision sequence for a named transaction in the last ninety days. An audit trail that captures only model decisions cannot answer the most common audit question, which is whether a human overrode the model and on what basis. Both fixes are technical, but both are first commissioned by an operator who has set the audit retention and retrieval expectation.
Chapter 5: Cross-cutting capabilities
The 4Ps describes four dimensions. Some practices sit across all four. They are gathered here so a reader cannot miss them while reading any single P in isolation. Each capability is mapped back to the Ps it serves, and each draws on a specific element of the Oxford AI Governance programme, the Wharton AI for Business Specialization, or the conformity assessment lineage that has become the operating standard in European AI governance work.
Testing, red teaming, and adversarial validation. Every material model entering production goes through three classes of pre-deployment testing. Performance testing checks the model does what is claimed on representative inputs. Robustness testing checks the model under distributional shift, adversarial inputs, and prompt injection where applicable. Red teaming, conducted by a team with a written brief to find the failure modes the build team did not consider, produces the third class. Red teaming is mandatory for any use case classified as high-risk under the EU AI Act, and good practice for everything else. Findings feed the model card. Where it lives: P3 model lifecycle (the testing pipeline) and P4 risk inventory (the residual risks after testing). Source: Oxford AI Governance; NIST AI RMF Measure function.
Bias, fairness, and explainability. Bias and fairness are governed by principle in P2 and verified by test in P3 and P4. A principle without a test is theatre. The testing pattern is established: select the protected groups relevant to the use case, define the fairness measure (demographic parity, equalised odds, calibration, or another defensible choice), measure pre-deployment, measure on a recurring cadence in production. Explainability is shaped by use-case risk class. A model recommending content can be explained at the cohort level. A model affecting an individual's credit, employment, or health needs explanation at the decision level, in language a non-technical reader can act on. Where it lives: P2 ethics review (decision on standards), P3 model lifecycle (testing), P4 monitoring (recurring evidence). Source: Wharton AI Applications in People Management; Oxford AI Governance; NIST AI RMF.
Continuous monitoring and drift detection. No model is governed only at deployment. Performance drift, data drift, concept drift, and population drift each erode model usefulness over time. The team operates a continuous monitoring dashboard for material models, tracking the model's headline performance metric, its fairness metric where applicable, and its drift indicators. Thresholds are set in advance, not negotiated after a breach. A breach triggers the incident playbook in P4, not a meeting to decide whether to act. Where it lives: P3 model lifecycle (the monitoring practice) and P4 risk inventory and incident response (the response to breaches). Source: NIST AI RMF Manage function; Oxford AI Governance.
Algorithmic Impact Assessment (AIA). An AIA is to AI what a Data Protection Impact Assessment is to personal data. For any use case classified as high-risk, the team produces a structured assessment covering the use-case description, the affected stakeholders, the rights and interests at stake, the potential harms by severity and likelihood, the mitigations in place, and the residual risk accepted. The AIA is signed by an executive accountable for the use case and reviewed at the operating cadence. Where it lives: P2 ethics review (the assessment as governance instrument) and P4 risk inventory (the residual risk record). Source: Oxford AI Governance; EU AI Act conformity assessment lineage.
Stakeholder engagement. Governance that omits the people affected by a model is governance built for the regulator, not the user. The framework requires that high-risk and externally facing use cases include a defined engagement step: identify the affected parties, gather their input through a fit channel, document what was heard, document what changed in response, document what did not change and why. Engagement is not a referendum. It is evidence. Where it lives: P2 ethics review (the engagement step) and P4 audit trail (the documented evidence). Source: Oxford AI Governance ethics module; OECD AI Principles.
Privacy and data protection. AI work intersects with privacy law everywhere. UK GDPR, EU GDPR, and the equivalent regimes in other jurisdictions are not optional. The framework requires that every use case has a privacy posture: lawful basis confirmed, data minimisation applied, retention limits set, cross-border transfer rules respected, automated-decision-making rights honoured. Privacy is not a separate workstream from AI governance. It is part of the same fabric. Where it lives: P1 data readiness (data fitness includes legal availability), P2 stated principles (privacy is one of the five recommended principles), P4 regulatory alignment (privacy law is part of the mapping). Source: Oxford AI Governance; UK GDPR; EU GDPR.
Sandbox and controlled testing environments. No new model goes into a production environment without time in a sandbox. The sandbox is a controlled environment with representative data, instrumented for the metrics that matter, and isolated from production systems. Controlled rollout, A/B testing, and shadow deployment are the techniques that move a model from sandbox to production at acceptable risk. Where it lives: P1 infrastructure readiness (the sandbox exists), P3 model lifecycle (the sandbox is a stage). Source: Wharton AI Strategy and Governance; NIST AI RMF.
Conformity assessment as a recurring ritual. The Oxford programme is grounded in the conformity assessment procedure for AI systems that has matured into the operating practice behind several EU-aligned governance platforms. The 4Ps adopts the same logic in operator form. The diagnostic in Part VIII is a self-conformity assessment, scored quarterly, validated by an external check. The model card, the AIA, the risk inventory, and the audit trail are the artefacts the conformity assessment relies on. The 4Ps does not require a specific vendor or platform. It requires the artefacts. The artefacts hold up regardless of who produces them. Where it lives: across all four Ps, expressed as the quarterly recalibration of the diagnostic. Source: Oxford AI Governance; capAI lineage; EU AI Act.
Supply resilience. AI capability is bought as a service but delivered from hardware the vendor does not fully control. Compute is a supply-constrained input, not an elastic utility, and when it runs short the vendor rations usage through rate limits, capacity tiers, and throttling. Because the scarcity is shared across the industry it is correlated: when demand spikes, providers throttle at the same time, so holding a second vendor is partial cover, not real cover. The framework requires four things for any use case that must not fail. Tier the use cases by what a supply interruption would cost. Hold a fallback the organisation controls, typically a smaller local or open model that degrades gracefully rather than stopping. Assess the vendor's deliverability rather than only its features: committed throughput or capacity reservation against best-effort terms, diversification across silicon, cloud, and power, the priority tier the contract buys, and the rationing track record. Then rehearse the day the tokens stop, as a named incident rather than a surprise. Where it lives: P1 infrastructure readiness (the owned fallback and supply headroom), P3 vendor and partner discipline (deliverability due diligence and contracts), and P4 risk inventory and incident response (capacity as a named risk and a rehearsed scenario). Source: operational resilience practice, the DORA-style discipline of critical-third-party and continuity governance, applied to compute.
The cross-cutting capabilities are not optional additions. They are the practices that distinguish a framework that satisfies the regulator on paper from a framework that produces governance that holds when challenged and creates business value.

Every one of them already lives inside the 4Ps. They are surfaced here because a reader who reads only one P should still encounter them.
Part V: How the 4Ps and 6Ps work together
The 6Ps from Team Performance Uplift and the 4Ps from this framework are not competing models. They are designed to work together. The 6Ps is the operating model for the team. The 4Ps is the governance frame around the team's AI use.
The 6Ps - People, Purpose, Principles, Practices, Productivity, and Performance - lives inside P3 (Practised) of the 4Ps. Specifically, it sits as the Team capability sub-dimension. A team that is doing the 6Ps well is producing the conditions inside which the 4Ps' Practised dimension can succeed.

The integration matters because most AI failures are not technical failures. They are team failures. A team without strong People practice cannot absorb the change. A team without clear Purpose deploys AI on the wrong problem. A team without Principles drifts at the moments operating discipline matters most. A team without lived Practices turns AI into accidental output. A team without honest Productivity measurement cannot tell where the waste sits. A team without honest Performance measurement cannot tell whether the AI is helping.
The 4Ps adds the four governance dimensions on top of this. Primed answers whether the team can use AI at all. Principled answers whether the team's AI use is ethically anchored. Practised, with the 6Ps inside it, answers whether the team's AI use is operationally disciplined. Protected answers whether the team's AI use is safe in front of customers, regulators, and auditors.
A team can be strong on the 6Ps and weak on the 4Ps. That team will perform well in the short term and accumulate governance debt. A team can be strong on the 4Ps and weak on the 6Ps. That team will be compliant and slow. Both are sub-optimal. The combination is the goal.
In practical terms, leaders adopting both frameworks should plan a single rollout. The 30-day rollout described in Part VI assumes the 6Ps is either in place or being adopted in parallel. Where it is not, the rollout takes longer because the team capability sub-dimension in P3 cannot be quickly created.
The mapping between the two frameworks is precise. Each of the six dimensions in the 6Ps maps to a specific operating outcome that supports the 4Ps' Practised dimension.
People supports P3 by making sure the team has the roles it needs. AI work demands a particular shape of team. Someone with model-level depth. Someone with use-case-level depth. Someone with product depth. Someone with risk depth. A working team has these roles either filled or accessible through a partner. The People dimension of the 6Ps ensures the roles exist and the development paths are real.
Purpose supports P3 by making sure the AI work is pointed at the right problem. The single largest cause of stalled AI use cases is unclear Purpose. The use case was technically interesting and commercially marginal. The 6Ps Purpose discipline forces a clear statement of why the work matters, to whom, and at what value.
Principles supports P3 by holding the team to known good operating patterns. Lean, agile, and the proven team dynamics behind both. AI work is novel enough that teams reach for new patterns and quietly lose the disciplines that produce quality at scale. The 6Ps Principles dimension keeps the team operating to patterns that have already been tested, so the operating discipline of P3 has a stable substrate to sit on.
Practices supports P3 by turning principles into the work itself. Where Principles defines what the team holds itself to, Practices is how it actually shows up: the standups, the retrospectives, the written decisions, the paired work, the review cadences. The 6Ps Practices dimension keeps the team doing the things that produce quality at scale rather than relying on heroics. Without lived Practices, the Use-case lifecycle and Operating cadence in P3 are paper artefacts.
Productivity supports P3 by measuring delivery and waste honestly. AI use cases consume time, money, attention, and compute. The team needs to know what it is shipping per unit of input, where the waste sits, and where the rework comes from. The 6Ps Productivity dimension makes that measurement routine so improvement is data-led rather than vibe-led. Productivity is the precondition for any Performance claim.
Performance supports P3 by giving the team honest outcome measurement. AI use cases require measurement at three layers: model performance, business outcome, and team performance. The 6Ps Performance discipline measures all three, separately, and does not allow one to be used as a proxy for another.
A worked example of the integration. A team is using the 4Ps and has scored Amber on P3 Practised. The diagnostic surfaces that the team capability sub-dimension is the lowest contributor. The investigation reveals that Performance measurement is the weakest of the 6Ps in that team. The intervention is to improve Performance, which lifts the team capability sub-dimension of P3, which lifts the overall P3 score. The framework operates as a single system. A weakness anywhere shows up somewhere.
The opposite is also true. A team strong on the 6Ps will find P3 lifts naturally as the 6Ps practice deepens. A team can use the 6Ps as the engine for improvement and the 4Ps as the measure of governance. The two frameworks reinforce each other.
Part VI: Usage instructions - a 30-day rollout

The 4Ps can be deployed in a team in thirty days. The rollout assumes a team of between ten and a hundred people, an executive sponsor, and a working week's commitment from a small core group. It does not assume a dedicated risk function or a large budget.
Week 1: Baseline assessment.
The first week is diagnostic. The core group works through the diagnostic checklists in Part IV for each of the four Ps and produces a scoring rubric based on Part VIII. The output is a single-page baseline showing where the team currently sits across the four dimensions. The baseline is shared with the executive sponsor and, ideally, with the team itself, so improvement is collective rather than imposed.
The baseline conversation surfaces the priorities. A team that scores Red on Primed has different first moves from a team that scores Red on Protected. The framework deliberately produces a sequence, but the rollout is shaped by where the team is starting.
Week 2: Design - principles and RACI.
The second week is design work. The core group drafts the team's AI principles, agrees them with the executive sponsor, and publishes them. In parallel, the RACI for material AI decisions is drafted and circulated. The first meeting of the ethics review process is scheduled.
Two outputs leave week two. The principles document, signed by the executive sponsor. The decision-rights RACI, published to the team. Both are short. Both are real.
Week 3: Wire-in - model lifecycle and risk register.
The third week is operational design. The core group drafts the model lifecycle and the use-case lifecycle, identifies the top three AI use cases currently in flight, and produces a model card for each material model already in production. The risk inventory is opened, populated with the current known risks, and assigned owners.
By the end of week three the team has principles, decision rights, lifecycles, model cards, and a live risk register. None of these is perfect. All of them are present and in use.
Week 4: Go-live and monthly cadence.
The fourth week is operationalisation. The operating cadence is published - weekly operational, monthly performance, quarterly portfolio. The first weekly operational review is held. The incident response playbook is drafted and a tabletop exercise is scheduled for the second month. The audit trail design is signed off.
By the end of week four the 4Ps is in operation. The team has a baseline, a set of principles, a RACI, lifecycles, model cards, a risk register, a cadence, and a playbook. The maturity is not at the target yet. The operating model is.
Months two through six.
The framework moves from rollout to operation. Each month, the team progresses sub-dimensions from Red to Amber, and from Amber to Green, in the diagnostic and scoring rubric. The expectation is that within six months the team is at or above Amber across all sub-dimensions, with at least P2 (Principled) and P4 (Protected) at Green. Within twelve months, all four dimensions should be at Green for a team operating in material AI use cases.
Month two is incident response rehearsal and vendor scorecard. The tabletop exercise scheduled in week four is run. Lessons are captured. The vendor scorecard is published for the top three vendors, with technical due diligence, data handling, performance commitments, indemnity, exit, and alignment with principles each scored.
Month three is workforce literacy and data readiness. The role-shaped literacy curriculum is launched for the two highest-volume roles. The data readiness assessment for the priority use cases is completed and gaps are scoped.
Month four is model lifecycle deepening. Model cards are produced for all material production models. Retraining or refresh schedules are agreed. The retirement protocol is exercised on at least one model.
Month five is regulatory mapping. The team's AI work is mapped against the EU AI Act, NIST AI RMF, and ISO 42001. Gaps are identified. Remediation is scheduled. The mapping is added to the audit pack.
Month six is the first quarterly portfolio review. The 4Ps baseline is rescored. The improvement plan for the next quarter is agreed. The board pack includes the heat map, the rescore, the incident log, and the plan.
The framework is not a once-and-done deployment. The operating cadence is what keeps it working.
The steady state.
After the first six months, the 4Ps becomes part of how the team operates. The weekly operational review handles live decisions on use cases and models. The monthly performance review handles trends, incidents, and vendor performance. The quarterly portfolio review handles strategic direction, rescore, and reprioritisation. The annual review handles the framework itself: whether the principles still fit, whether the rubric still discriminates, whether the operating cadence still produces decisions.
The team will know the framework is working when three things are true. Decisions are made faster, with less rework. Incidents are smaller and shorter. Customers, regulators, and auditors find the team's evidence current and clear. None of these is a marketing claim. Each is observable in the team's calendar and decision log.
A worked rollout: a fictional but representative team.
To make the rollout concrete, consider a fictional team called Atlas Health, a 180-person digital health business with a recently funded Series B and two AI use cases in flight. Use case one is a clinician-facing summary generator that produces a short narrative of a patient visit. Use case two is a customer support copilot that drafts replies to inbound queries.
Week one. The core group is the Head of Product, the Head of Data, the Chief Medical Officer, and the General Counsel. The diagnostic produces a 4Ps baseline. Atlas Health scores Amber on Primed (leadership literacy is high, workforce literacy is uneven), Red on Principled (no written principles, RACI exists only for the clinician-facing use case), Amber on Practised (use-case lifecycle exists informally, model lifecycle is undocumented), and Red on Protected (no risk register, no model cards, audit trail is technically present but not retrievable inside an hour).
Week two. Principles are drafted. Five principles emerge: clinical safety, patient privacy, transparency to clinicians, human oversight on clinical decisions, and fairness across patient demographics. The principles are signed by the CEO and circulated to the team. The RACI for material AI decisions is drafted. Three decisions are defined: model selection, deployment to production, and retirement. The Chief Medical Officer is Accountable for all three. Responsibility is split between the Head of Data (model selection), the Head of Product (deployment), and the Head of Engineering (retirement).
Week three. The use-case lifecycle is documented: idea, scoping, build, pilot, scale, monitor, retire. The model lifecycle is documented. Model cards are produced for the two production models. The risk inventory is opened. The clinician-facing summary generator carries six risks: hallucination in clinical context, omission of relevant detail, demographic bias, data residency, vendor model upgrade, and clinician over-reliance. The customer support copilot carries four risks. Each risk has an owner and a current mitigation.
Week four. The operating cadence is published. Weekly operational review on Tuesday morning. Monthly performance review on the second Wednesday. Quarterly portfolio review in the last week of each quarter. The first weekly review is held and produces three decisions, all recorded in a decision log. The incident response playbook is drafted. A tabletop exercise is scheduled for week eight.
Month two. The tabletop exercise is run with a scenario: a clinician reports that the summary generator omitted an allergy. The playbook is exercised. Two gaps are identified: there is no defined hand-back protocol from product to engineering when a clinical issue is reported, and the audit trail does not capture which clinical record the summary was generated from. Both are added to the improvement plan.
Month three. The role-shaped literacy curriculum is launched. Clinicians get a two-hour module on the summary generator's strengths and weaknesses with worked examples. Customer support agents get a one-hour module on the copilot's confidence signals and the human-in-the-loop rule. Data readiness for the third use case under consideration (a triage prioritisation model) is assessed and found to be insufficient. The use case is paused.
Month four. Model cards are updated. The summary generator's vendor announces an upgrade in eight weeks. The dependency map is reviewed. The upgrade window is added to the cadence. Retraining schedules are documented.
Month five. Regulatory mapping is completed against the EU AI Act, NIST AI RMF, and ISO 42001. The clinician-facing summary generator is classified as high-risk under the EU AI Act, which triggers a set of obligations. The team is six weeks ahead of the obligations because the 4Ps work has produced the artefacts. The mapping is added to the audit pack.
Month six. The first quarterly portfolio review is held. The 4Ps baseline is rescored. Atlas Health is now Amber on all four Ps, with P2 (Principled) on the cusp of Green. The improvement plan for the next quarter is agreed. The board pack contains the heat map, the rescore, the incident log, the regulatory mapping, and a one-page summary of the AI portfolio.
The example is fictional. The pattern is real. The thirty-day rollout produces operating discipline. The five months after produce maturity. The first quarterly portfolio review is the moment the framework becomes durable, because by then the cadence has run a full cycle and the team can see the rate of improvement.
Common rollout pitfalls.
Trying to do everything at once. The 4Ps is sequential by design. Teams that try to stand up all four Ps simultaneously usually produce partial versions of all four. Teams that follow the sequence usually produce working versions of each.
Skipping the baseline. The temptation is to start with the design work and skip the diagnostic. The diagnostic is what produces the priority list. Without it, the design work is uniformly applied and the priorities are missed.
Outsourcing the rollout. External help is useful for the diagnostic and the design. The rollout has to be carried by the team that will live with the outcome. A consultant who writes the principles will produce principles the team does not live by. A consultant who facilitates the team writing its own principles will produce principles that hold.
Treating the cadence as optional. The cadence is the framework. Without the operating cadence, the artefacts produced in weeks one through four become a folder no-one opens. The cadence is what keeps the artefacts current.
Confusing the rubric with the framework. The 1-5 scoring rubric is a measurement tool. It is not the work. A team that obsesses over the rubric and neglects the diagnostic conversations behind it will produce a number that does not reflect operating reality.
Part VII: Two audience summaries
The 4Ps applies to organisations of every shape. The implications differ. The two summaries below are written for the two audiences who most often ask for help: enterprise leaders and startup or SME founders. Each summary can be lifted into a profile, deck, or proposal.
For corporate and enterprise leaders
Position: Enterprise AI Transformation Architect and Advisor.
You are running an enterprise that has decided AI is strategic. You have one or more business units running pilots. You have a board that wants quarterly evidence of progress. You have a regulator paying attention, customers asking questions, and a workforce that is some combination of excited, anxious, and shadow-adopting tools you have not approved.
You need someone who can do four things at once. Shape AI strategy at the executive level. Implement AI-led change in operations. Stand up the governance the regulators expect. And do all of that with people who have been told different things by different consultants over the last two years.
My recent work answers each of those. At Haleon PLC I led the AI-enabled uplift of product, engineering, and data teams that produced a 45% performance lift, a 15% cost reduction, and roughly 99% delivery predictability across the engagement. The work coordinated EY, BCG, and Bain consultants under a single operating model, and included leading an AI-powered smart retail shelf solution with Bain. At GSK PLC the same template delivered a 40% lift on similar metrics. At Ogier Group, a legal and financial services firm, I designed and implemented corporate and data governance and risk management frameworks in a heavily regulated environment. At Teamsmiths the proprietary AI SaaS product has delivered a 45%+ uplift in team throughput and roughly 99% delivery predictability in three months using generative AI large language models, retrieval-augmented generation, and the proprietary uplift model.
The credentials behind the work are Oxford AI Governance at Saïd Business School (December 2025, 100% grade) and the Wharton AI for Business Specialization (March 2025, four certificates including AI Strategy and Governance at 100%). The earlier track record covers a £25M accounts receivable recovery at Philips Italy that three predecessors had failed to clear, a project management application at Philips that safeguarded £22M annually, EU data migrations at Reuters, strategy at IF Consulting that contributed to $3.2bn new revenue across Orange, VW, and ICI, a financial analytics product at Knight Ridder/McClatchy that generated £1m in year one, the strategic planning process and governance framework for a UK Government Executive Agency transition, infrastructure projects at Balfour Beatty between £20M and £25M, and the growth of a Gartner portfolio from $3M to $10M over five years as Senior Product Director, EMEA Digital Media and eCommerce, recognised with the Chairman's award.
The 4Ps AI Governance Framework is the operating frame I use with enterprise clients. It produces governance that satisfies the EU AI Act, NIST AI RMF, and ISO 42001 without slowing the team. It integrates with the 6Ps operating model from the Team Performance Uplift book to give you a single rollout that compounds.
The engagement model is straightforward. A six-week diagnostic produces a baseline across the 4Ps for one business unit. A ninety-day initial implementation delivers the operating model. A twelve-month programme delivers a Green rating across the framework. Each phase has its own commercial terms and clear exit. The implicit ask is straightforward too. Hire me to shape and implement profitable AI-led change with responsible governance.
A note on the multi-agent architecture experience. The five production agentic AI platforms are not demonstrations. They are running systems with end users, model providers, retrieval stacks, observability tooling, and the operational and commercial reality that comes with each. The architectural patterns that work at startup scale are the same patterns that work at enterprise scale once they have been hardened. Enterprise clients hiring for AI architecture work get the benefit of a practitioner who has implemented the patterns, not only read about them.
A note on the multilingual reach. Enterprise AI work in 2026 is not English-only. EU regulation is drafted in twenty-four official languages. Customer-facing AI applications need to operate in the language of the customer. Italian language certification from the Università per Stranieri, fluent business English, and conversational Yoruba together provide working coverage across the United Kingdom, Italy, and significant parts of West Africa. The point is not the languages themselves. The point is a practitioner who has thought about how the framework operates in a non-English-first environment, because they have lived in one.
A note on public engagement. Recent public speaking has covered AI Transformation and Agentic AI on the Agile Sherpa webinar, AI, Data and Energy Usage on a University of Sheffield panel with Forrester analysts, and AI-Powered Team Performance Uplift in a Surrey Chamber of Commerce keynote. The engagement record matters because the framework is intended to be teachable. A framework that cannot be explained in a thirty-minute keynote is not a framework. It is a deck.
For startup and SME founders
Position: Founder-operator with five production AI platforms shipped.
You are running a company with less money than you would like and more ideas than you have hands. You have heard the AI noise. You have started using ChatGPT, Claude, or both. You are wondering whether to build something proper, buy something proper, or wait. You do not have a Chief AI Officer. You will not have one for a while.
You need someone who has shipped. Not advised on shipping. Not theorised about shipping. Actually shipped, in production, with paying customers or live users.
My recent shipping looks like this. Five production agentic AI platforms built and operating: Revenue-Risk-Radar, Songita, Deputee, SDRAgent, AIMarketer, and Piccolingo. Each is a different bet on a different problem. Each runs in production. Each was built with the constraints you know: small team, no time, real users. The 4Ps AI Governance Framework was built to be deployable by a team that does not have a risk function.
Earlier work tells you what discipline looks like at small scale. Co-founding G2Guide from 2005 to 2020 produced partnerships with Reed Elsevier and IDOX and delivered up to 53% cost savings on supplier contracts for clients who needed the money saved more than they needed the report.
The lean credentials are the relevant ones for you. Oxford AI Governance gives you the regulator answer when you need it. Wharton AI for Business gives you the commercial frame. The five production platforms give you the engineer's pattern library for agentic AI. The multilingual reach matters if your buyers sit across the United Kingdom, Italy, or African markets. I work fluently in English and Italian, with conversational Yoruba, certified by the Università per Stranieri.
The 4Ps for a startup is a one-page plan, not a programme. Week one: pick one use case, name an owner. Week two: write five principles, agree the RACI, document the human-in-the-loop rule. Week three: stand up a lightweight risk register, write one model card. Week four: go live, agree the monthly cadence. The framework does not need a consulting team. It needs a founder who is willing to be precise for thirty days.
The implicit ask is straightforward. A partner who has built and shipped, not just advised. The engagement model is small and clear. A two-day workshop produces the plan. A ninety-day partnership delivers the first use case into production with the 4Ps wrapped around it. After that, you have an operating model you can run yourselves.
There is a second reason a founder might want this kind of partnership. Investors are starting to ask about AI governance the way they used to ask about cyber posture. A startup with a one-page 4Ps baseline, a written set of principles, and a model card for its production model is a startup that has already answered the diligence question. That is not a marketing exercise. It is a working artefact that takes two weeks to produce and a quarter to mature. The cost of producing it before an investor asks is a fraction of the cost of producing it during an investor's diligence.
There is a third reason. Founders who have shipped agentic AI know what happens when a model upgrade changes the answer the system gives. The same prompt produces a different output on Tuesday than it did on Monday. The model card, the operating cadence, and the audit trail are what let a founder explain that to a customer without losing the customer. The 4Ps is the thinnest viable wrapper around an agentic stack that still produces an answer to that conversation.
Part VIII: Diagnostic and scoring rubric

The 4Ps uses a 1-5 maturity score on each sub-dimension, totalling to a maximum of 100. Each P contributes 25 points (5 sub-dimensions x 5 maturity points). The maturity bands are coloured Red, Amber, and Green to support quick reading.
Maturity levels.
Level 1 (Red, score 1): Absent. The sub-dimension is not addressed. > There is no current intent to address it.
Level 2 (Red, score 2): Aware. The sub-dimension is recognised. > Discussion is ongoing. No artefact exists.
Level 3 (Amber, score 3): Defined. The sub-dimension is documented. > Operation is partial. The artefact exists but is not yet lived.
Level 4 (Amber, score 4): Operating. The sub-dimension is in > operation. Evidence is current within the last quarter.
Level 5 (Green, score 5): Embedded. The sub-dimension is operated to > a published cadence. Evidence is current within the last month. > Continuous improvement is visible.
Colour bands.
Red: total score for a P below 13. Action required this quarter.
Amber: total score for a P between 13 and 19. Improvement plan in > place.
Green: total score for a P 20 or above. Maintain and improve.

A team's overall 4Ps score is the sum of the four P scores. The thresholds for the overall score follow the same proportions.
Overall Red: below 52. Material risk. Executive attention required.
Overall Amber: between 52 and 76. Recoverable. Improvement plan and > executive sponsor required.
Overall Green: 77 or above. Operating well. Sustain through cadence.
The rubric is deliberately simple. A more elaborate scoring system was tested in early pilots and found to slow the diagnostic without improving its quality. Sixty minutes of conversation against the seven-question checklist in each chapter, anchored to the 1-5 scale and the band thresholds, produces a baseline that holds up.
The diagnostic should be scored by the operating team itself, validated by an external check, and recalibrated quarterly. Self-scoring without external check produces drift. External scoring without self-input produces resistance.
A worked example. A team scoring 4 on Leadership literacy, 3 on Workforce literacy, 3 on Data readiness, 3 on Infrastructure readiness, and 4 on Use-case clarity scores 17 on P1, which sits in the Amber band. The improvement plan would prioritise Workforce literacy, Data readiness, and Infrastructure readiness to lift each from 3 to 4 over the next quarter. The next quarterly recalibration would re-score against the same rubric.
Scoring discipline.
Three rules keep the scoring honest.
Evidence over assertion. A 4 (Operating) is awarded only when the team can produce current evidence within the last quarter. A 5 (Embedded) is awarded only when the team can produce current evidence within the last month and demonstrate continuous improvement. Assertion without artefact stays at 2 (Aware) or below.
Conservative when in doubt. If a sub-dimension sits between two levels, the score is the lower one. The next quarter's plan then includes the work to clear the higher level unambiguously.
Calibrate quarterly. A team's view of its own maturity drifts over time. The quarterly external check is what keeps the scoring honest. The external check does not need a third party. It can be done by a peer team or by the executive sponsor.
Reading the heat map.
The 4Ps heat map is a four-by-five grid. Four rows for the four Ps. Five columns for the five sub-dimensions of each. Each cell is coloured Red, Amber, or Green according to the score.
The diagnostic conversation reads the heat map in three passes.
The first pass looks at the rows. Which P has the lowest total? The answer points to the framework dimension needing most attention.
The second pass looks at the columns. Are particular sub-dimensions consistently weak across multiple Ps? For example, is the team consistently weak on the regulatory and audit-facing sub-dimensions, regardless of P? That points to a structural gap rather than a dimensional one.
The third pass looks at the cells. Where is the single weakest sub-dimension? That is the priority for the next month.
A team that does the three passes regularly will produce more useful conversations than a team that fixates on the total score.
Common rubric misuses.
Treating the rubric as a benchmarking tool against other organisations. The rubric is calibrated for a team's own progression. Cross-team comparison is possible only when the scoring discipline is identical, which is rare in practice.
Using the rubric to justify investment without addressing the work behind it. The score is a measurement, not a budget request. A Red on Infrastructure readiness might be solved by a vendor change, by a configuration change, or by retiring a use case. The rubric does not specify which.
Allowing the rubric to slip from quarterly to annual. The framework is designed for quarterly cadence. Annual rescoring loses the early warning signal and produces too large a jump between data points.
Part IX: Beyond the assessment - the product layer
The 4Ps is a framework. The framework can be operated with paper and discipline. Most teams will want a product layer that takes the friction out of the operation. The 4Ps Governance Cockpit is the product the framework supports, hosted on the public web and available to enterprise customers in a private deployment.
The Cockpit offers five capabilities.
Self-assessment. The team completes the diagnostic against each of the twenty sub-dimensions. The Cockpit produces the score, the band, the heat map, and the priority list. The output is a one-page baseline and a recommended ninety-day plan.
AI inventory. The team registers each AI use case in production or in flight. Each entry captures the owner, the data dependencies, the platform, the risk class, the human-in-the-loop rule, and the linked model card. The inventory is the source of truth for everything else.
Risk heat map. Risks are captured against each use case, classified by impact and likelihood, with owners and current mitigations. The heat map is the operating view used in the monthly cadence. It is exportable to internal audit and to the board.
Model card generator. The Cockpit produces a model card from a structured set of prompts. The card follows the convention recommended by the framework, is exportable to PDF and HTML, and is versioned. The generator reduces a model card from a half-day of drafting to under thirty minutes.
30-day rollout plan generator. Based on the baseline and the priority use cases, the Cockpit produces a personalised version of the rollout plan in Part VI. The plan is owned and editable. It feeds the operating cadence.
The Cockpit is built to the principles it supports. The AI it uses is disclosed. The data the customer enters belongs to the customer and is not used to train models. The product has a documented model lifecycle of its own, model cards for its own models, and an audit trail accessible to administrators.
Pricing is positioned for two audiences. A free tier supports an individual leader running the diagnostic for a single team. A paid tier supports a team of fewer than a hundred. An enterprise tier supports a private deployment with single sign-on, audit-trail export to the customer's systems, and a customer-controlled retention policy.
The product layer is the difference between a framework people admire and a framework people use.
Roadmap for the product layer.
The Cockpit roadmap is sequenced to the maturity of the customer base, not to the maturity of the technology. Early customers need the diagnostic and the rollout plan. Mid-maturity customers need the inventory and the model card generator. High-maturity customers need the heat map, the cadence support, and the audit-trail export.
Phase one delivers the diagnostic, the AI inventory, and the model card generator. These are the artefacts every customer needs from the first month.
Phase two delivers the risk heat map, the rollout plan generator, and the cadence support. These artefacts are activated as the customer's operating model matures.
Phase three delivers the audit-trail export, the regulatory mapping helpers, and the integration with the customer's existing GRC tools. These artefacts make the Cockpit usable by enterprise audit and compliance functions.
Phase four delivers the partner network. Trained partners deliver the diagnostic and the rollout in customer environments using the Cockpit as the shared workspace. The partner network is what lets the framework reach scale without compromising quality.
Each phase is delivered against the same principles the Cockpit supports. The product is a working demonstration of the framework, not a brochure for it.
Part X: Governing the self-improving company
A new kind of company has appeared. It is built around closed AI loops that sense, decide, act, and learn, with the artefacts of every action recorded back into a shared brain the team and the AI both work from. The pattern was crystallised in a 2026 Y Combinator talk by Tom Blomfield as "the self-improving company". The idea is sound and the form factor will spread. The 4Ps governs this case, and governs it better than any framework built for a static system. This Part explains how.
Two senses of self-improving, kept apart. Before any framework can be applied, two ideas have to be separated. The first is the organisational pattern Blomfield describes: a company designed as a set of closed loops that compound because the loops capture their own evidence. The components are operational. Sensors and data, a policy layer, a tool layer, quality gates, a learning mechanism. Humans sit at the edges where relationships, judgment and trust live. The work that happens here is the work of any disciplined team, made legible to AI. The second sense is recursive self-improvement in the AI-safety sense: an AI system iteratively modifying its own architecture, weights, or training to become more capable, with each round making the next easier. Almost no enterprise is doing this in any operational sense today, and the recent literature is converging on the view that even when an AI proposes a better training recipe, the loop remains gated by compute, supply, eval, and organisational judgment. Both senses belong in a governance framework, but the first is the live, mass-market case and the second is a profound risk class that needs naming when it appears. Treating them as one over-governs the first and under-governs the second.
Where the 4Ps fits the self-improving company. Blomfield's components map almost one to one onto the four Ps. Sensors and data live in Primed: literacy, data readiness, infrastructure readiness, use-case clarity. The policy layer is Principled: stated principles, decision rights and RACI, ethics review, transparency, human-in-the-loop. The tool layer and the quality gates are Practised: the use-case lifecycle, the model lifecycle, operating cadence, vendor discipline, and the 6Ps team capability that operates them. The learning mechanism is Protected: audit trail, drift detection, incident response, conformity assessment as a recurring ritual. The cross-cutting capabilities of Part IV Chapter 5 are exactly the practices a loop-based company needs in addition: continuous monitoring, Algorithmic Impact Assessments for high-risk loops, stakeholder engagement, privacy posture, sandboxes for change, and the supply resilience that keeps the loops running when compute is rationed.

The harness language carries through. A self-improving loop is the rawest power AI offers a business because it compounds. Left loose it bolts in a specific and recognised way, which the third sharpening below names. Harnessed by the 4Ps it pulls more than any one-off model deployment will. Good governance turns a loop into a compounding asset.
Sharpening one. Treat the loop as a governable artefact in its own right. A self-improving company runs dozens of loops. Treating loops as instances of use-cases is a category error because a loop has its own change rhythm, its own authority, its own failure modes. The 4Ps adds a Loop Lifecycle inside Practised, alongside the existing use-case and model lifecycles. Each loop, at the moment it is proposed, names six things. An owner who is accountable for the loop, not just the use case. An authority class. A feedback channel. A kill switch and the conditions that trigger it. A published review cadence. A retirement criterion. The authority class is the centre of gravity. Most loops should not be autonomous, and the framework makes that the default.

The classes climb from read-only, where the loop only observes and reports, through recommend, where the loop suggests and a human acts, to act-with-reversal, where the loop acts but a human can reverse inside a window, and finally to autonomous-with-audit, where the loop acts on its own and humans review on cadence. Governance cost climbs with the class. Read-only loops need transparency and source citation. Recommend loops need explainability and an override log. Act-with-reversal loops need a reversal SLA, an audit trail, and a named owner. Autonomous loops need a kill switch, an Algorithmic Impact Assessment, a sandbox change discipline, and a monthly human review. Every loop names its class on day one and the class drives the controls.
Sharpening two. Name feedback integrity as a first-rank risk class. Inside Protected, the framework already covers model risk and supply risk. Self-improving loops introduce a third risk class that few enterprise frameworks have named yet: feedback collapse. When a loop's output is recycled as its input, quality degrades in a slow, invisible way. Each cycle drifts. Biases compound. The team often does not notice for weeks. The current literature, including Epoch AI's analysis of when fresh human text runs short, puts this as an enterprise risk worth naming from 2026 through 2032. The controls are mundane and powerful.

Three disciplines hold the line. An input ledger names every input source the loop trains on and flags AI-generated content. A clean evaluation set is held in escrow and is never touched by loop output. Drift monitoring runs continuously against the clean set and triggers an alert when divergence crosses a defined threshold. Feedback integrity then earns a place in the risk inventory, in the model cards as a named failure mode, and in the incident response playbook as a rehearsed drill. A team that has run the feedback-collapse drill once will spot the next collapse in days rather than weeks.
Sharpening three. Make the surveillance guardrail explicit. The Blomfield idea has a sharp human edge. A company brain that records every action and every artefact is, by construction, one inch from worker surveillance. The harness rule already in the framework, harness the AI not the people, becomes a hard governance commitment in a self-improving company rather than a stylistic preference. Two things follow. Principled adds a transparency standard that names what the company records, why, who can see it, and what protections are guaranteed. Protected adds a labour-law line to the regulatory mapping: GDPR Article 88 in Europe, the relevant works-council rights in EU jurisdictions, and worker-monitoring statutes elsewhere. Legibility serves the work, not the worker. A loop that captures evidence about how the work happens is governable. A loop that captures evidence about the worker requires a different conversation, a different process, and a different consent.
The opportunity, named directly. A self-improving company governed by the 4Ps captures the compounding upside that loops produce while keeping the risks of feedback collapse, drift, and surveillance off the floor. The mechanism is not subtle. A loop running at the read-only class for a quarter, escalated to recommend once the team trusts the evidence, escalated again to act-with-reversal once the reversal SLA is met, builds a compounding asset whose value rises faster than its cost. The framework's value-creation frame is at its most literal here. Compliance held, business value created, every cycle.
The 30-day rollout, adapted for the self-improving company. The four-week pattern in Part VI carries through with one substitution. Week 1 is the baseline, including a loop inventory: every existing loop named, owned, and classed. Week 2 is the design, including the loop authority taxonomy and the input-ledger discipline. Week 3 is the wire-in, with at least one loop instrumented for drift monitoring against a clean evaluation set. Week 4 is the go-live, with the kill switches tested, the labour-law line added to the regulatory mapping, and the first loop retirement criterion published. The cadence after the rollout is the same as for any 4Ps engagement, with one addition: the monthly performance review surfaces the drift report on every loop, and the quarterly portfolio review rebalances the loop authority classes based on evidence rather than enthusiasm.
The summary line. The self-improving company is the strongest application the 4Ps has, not because the framework was written for it but because the framework's premise applies cleanly: operator-first governance that creates business value, harnessing power rather than caging it. Three sharpenings make the framework fit the case in full: the loop as an artefact, feedback integrity as a risk, and the surveillance guardrail as a line. Add them and the framework governs the compounding company. Skip them and the loops eventually bolt.
Appendix A: Glossary
Agentic AI. AI systems that operate over multiple steps with some degree of autonomy, often combining language models with tools, memory, and orchestration.
Audit trail. A retrievable record of the decisions, inputs, outputs, and human interventions related to an AI system, used for accountability and incident review.
Decision rights. The codified assignment of who is Responsible, Accountable, Consulted, and Informed for each material decision related to an AI system.
Ethics review. A structured forum at which a proposed or in-flight AI use case is examined for fairness, harm, transparency, and proportionality, with documented outcomes.
EU AI Act. The European Union's regulation establishing harmonised rules on artificial intelligence, including a risk-based classification of AI systems and obligations for providers and deployers.
Foundation model. A large pre-trained model that can be adapted to a wide range of downstream tasks.
Generative AI. AI systems that produce content - text, image, code, audio, video - rather than only classifying or scoring inputs.
Human-in-the-loop (HITL). An operating pattern in which a human reviews, approves, or otherwise intervenes in the output of an AI system before the output is acted on.
ISO 42001. The international standard specifying requirements for an artificial intelligence management system within an organisation.
Large language model (LLM). A foundation model trained primarily on text data, capable of producing language output in response to prompts.
Maturity rubric. The 1-5 scoring instrument used to assess each sub-dimension of the 4Ps framework, with colour bands of Red, Amber, and Green for quick reading.
Model card. A short, structured document describing a model's purpose, training data, performance, limitations, intended use, and accountable owners.
NIST AI RMF. The National Institute of Standards and Technology's AI Risk Management Framework, a voluntary framework for managing risks associated with AI.
OECD AI Principles. A set of values-based principles for trustworthy AI adopted by member countries of the Organisation for Economic Co-operation and Development.
Operating cadence. The published rhythm of weekly, monthly, and quarterly reviews through which the organisation operates its AI work, with named attendance and recorded decisions.
RACI. A decision-rights tool identifying who is Responsible, Accountable, Consulted, and Informed for a given decision or activity.
Retrieval-augmented generation (RAG). A technique that combines a language model with an external retrieval step, allowing the model to ground its output in specified sources.
Risk class. A classification of an AI use case by the impact and likelihood of harm, used to determine the level of governance and human oversight required.
Risk inventory. A maintained list of AI-related risks specific to the organisation's use cases, with classification by impact and likelihood, named owners, and current mitigations.
Shadow AI. Unsanctioned use of AI tools by employees outside the organisation's official platform and policies. A high level of shadow AI indicates a gap between the official tools and the work the team is trying to do.
Tabletop exercise. A rehearsed walkthrough of an incident scenario with the named response team, used to test the playbook before a real incident.
Use-case clarity. The discipline of describing each AI use case in terms of business outcome, owner, success measure, and risk class, before any platform or model commitment is made.
Use-case lifecycle. The defined stages a candidate AI use case moves through, from idea to retirement, with entry and exit criteria at each stage.
Vector store. A database optimised for storing and retrieving embeddings, often used in retrieval-augmented generation pipelines.
Vendor scorecard. A structured assessment of an AI vendor against the organisation's standards, covering technical due diligence, data handling, indemnity, performance commitments, exit terms, and principle alignment, updated at least quarterly.
4Ps Wheel. The diagrammatic representation of the framework as four equal quadrants rotating continuously around the team at the centre.
4Ps Governance Cockpit. The product layer hosted on the public web that supports the framework through self-assessment, AI inventory, risk heat map, model card generator, and rollout plan generator.
Appendix B: References
The 4Ps draws on the following bodies of work. The list is for the reader who wishes to deepen any dimension. No citations are fabricated. Where a URL is given it is the canonical source.
Saïd Business School, University of Oxford. AI Governance programme. > (Author certification, December 2025, 100% grade.)
The Wharton School, University of Pennsylvania. AI for Business > Specialization, comprising AI Strategy and Governance (100% > grade), AI Fundamentals for Non-Data Scientists, AI Applications > in Marketing and Finance, and AI Applications in People > Management. (Author certification, March 2025.)
European Union. Regulation on Artificial Intelligence (the EU AI > Act). Official Journal of the European Union.
National Institute of Standards and Technology, United States > Department of Commerce. AI Risk Management Framework (AI RMF 1.0).
International Organization for Standardization. ISO/IEC 42001 > Information technology - Artificial intelligence - Management > system.
Organisation for Economic Co-operation and Development. OECD > Principles on Artificial Intelligence.
Osu, S. Team Performance Uplift. The 6Ps operating model for team > performance. (Companion volume.)
Saïd Business School, University of Oxford. Blockchain Strategy > programme. (Author certification.)
University of Westminster. B.Eng (Hons). (Author qualification.)
The Management School, University of Sheffield. MBA. (Author > qualification.)
University of Colorado. Agile Leadership programme. (Author > certification.)
Scaled Agile. Leading SAFe certification. (Author certification.)
Scrum.Org. Professional Scrum Master (PSM). (Author certification.)
ICAgile. ICP-ACC, Agile Coach. (Author certification.)
Università per Stranieri. Italian language certification. (Author > certification.)
School of Practical Philosophy. Practical Philosophy training. > (Author study.)
Appendix C: Diagrams
The framework diagrams now appear inline where they are discussed. The 4Ps wheel and the anatomy of the harness are shown in Part III, the framework in one page. A 30-day rollout process flow is maintained as a companion file (4Ps_ProcessFlow.svg) and can be embedded on request. The descriptions below summarise each diagram for reference.
[Diagram: The 4Ps Wheel. Four equal quadrants - Primed, Principled, Practised, Protected - arranged around a central hub showing the team. The wheel rotates clockwise, indicating continuous operation. The five sub-dimensions of each P are arranged radially within their quadrant.]
[Diagram: The 30-day rollout process flow. Week 1 Baseline assessment. Week 2 Design (principles and RACI). Week 3 Wire-in (lifecycles and risk register). Week 4 Go-live and monthly cadence. The flow then enters the steady-state operating cadence loop: weekly operational review, monthly performance review, quarterly portfolio review.]
Closing note
A framework is a tool. A tool is useful only when picked up. The 4Ps will help a team that uses it and not help a team that prints it.
The reader who has finished this document now has four choices. Put the framework down and continue as before. Adopt a single P, the one most relevant this quarter. Run the 30-day rollout for one team. Run the rollout for the organisation.
Each choice has consequences. The first is honest if the team is not yet ready. The second is the most common starting point in practice. The third is the most common starting point for a serious engagement. The fourth is the right choice for an organisation in which AI is already material to the business.
The framework will continue to evolve. The early adopters who give feedback shape the next version. The product layer described in Part IX is the channel through which the feedback is collected and the framework refined. The author can be reached at segun.osu@teamsmiths.com.
Responsible AI is not a destination. It is a discipline. The 4Ps is a working name for that discipline. Use it well.
End of document.
