Overview
What this was
The development cycle translated stakeholder perception into observable behavior, embedded that behavior into normal management routines, and returned six months later to test whether the experience of working with the leader had changed.
- Designed the leadership framework and the 360 instrument built against it
- Facilitated the senior-leader team effectiveness sessions
- Ran every individual debrief and built the 90-day plans with participants
- Moved verification into the participant's own business review
- Built and administered the six-month reassessment
The system
Diagnosis, build, and how it ran
The constraint
The leaders already had plenty of feedback. What they lacked was a mechanism that carried insight into the quarter. Peers kept building shadow trackers. Direct reports kept re-deriving their own priorities each week.
So I built the cycle as one system with a single test at the end: does the experience of working with this leader change, and does the change survive after the coach steps away?
Phase 1 · Diagnose
The gap between intent and experience
A 360 usually arrives as a score. I use it to find where a leader's intent and their stakeholders' experience diverge, which patterns repeat across groups, and which few behaviors cause the most friction.
What the diagnosis surfaced
I read the report for divergence: where a leader's own rating and their stakeholders' ratings separate, which groups see it, and which few behaviors sit underneath the gap. The gap identifies where intent and experience differ.
What made it usable
- Three lowest items named: one on expectations, one on priority stability, one on visible closure
- One strength selected deliberately to help carry the change
- One real gap, psychological safety, explicitly deferred: three priorities at once produce three partial results
Define
Vague development goal → observable operating behavior
This is the pivot the whole system depends on. Every commitment is written so a third party can tell whether it happened without asking the participant.
Written so it can be observed
Issue any priority change in writing within 24 hours and name what is being dropped to make room.
Written so it can be observed
Write a one-page role scorecard for each direct report (outcomes owned, decision rights held, criteria their work is judged against) and review it live with each person.
Written so it can be observed
Post closure in the same thread where the commitment was made, on or before the date; if the date will slip, post a revised date before the original passes.
Written so it can be observed
Send peer function leads a monthly written recap: what closed, what slipped, what changed.
Phase 2 · Practice
Facilitating senior leaders against evidence
The team effectiveness session for 20 executives, VPs, directors, and functional heads tests leadership identity claims against observable behavior and ends in a commitment written for third-party verification.
I use the five team conditions as diagnostic lenses, asking leaders to test whether their teams actually experience the conditions they believe they create.
Practice
Five conditions, five diagnostic questions
Select a condition to see the question I put to the room.
People can raise risk, admit mistakes, and challenge decisions without managing the leader's reaction.
“When was the last time someone publicly challenged one of your decisions?”
Holding the tension
When a senior leader says “people here can speak up, we're already an open culture,” I treat it as a hypothesis to test against the team's experience. Then I ask the room the question above and stop talking for eight to ten seconds.
If the answer is hard to retrieve, that is data. The session closes with each leader reading their committed behavior out loud to a partner. The partner tests one thing: could they recognize the behavior when it happens? Anything too vague gets rewritten.
Facilitation excerpt uses fictional scenarios and participant details. The original program deck remains Company Confidential.

What the session produced
The Project Aristotle senior-leader workshop is measured through documented follow-through after the session.
- 20 participants
- 19 documented a behavioral commitment
- 16 completed the 30-day follow-up
- 14 could point to a concrete application of the behavior in their work
Counts come from the session record and 30-day participant follow-up. They reflect self-reported follow-through.
Phase 3 · Turn insight into behavior
Feedback → behavior → evidence
Two priorities, six commitments, each with a cadence, an evidence source, and a measure. The evidence is captured where the work already happens.
Define
How a commitment is constructed
Select a step to see what it produces.
The checkpoints did real work
At 30 days the plan was at risk. Several role scorecards were still unwritten, and the missing ones belonged to the regions in the middle of a structure change, deferred on the reasoning that the roles might move.
The adjustment was to write them against the current role, dated, marked provisional. Clarity work gets postponed precisely when conditions are unclear, which is when it matters most. By 60 days the practices were holding without prompting and support moved from weekly to biweekly.
Phase 4 · Scale
Designed to run inside the business
Four mechanisms move reinforcement out of my calendar and into the business.
1. Manager-led verification
Traditional
L&D designs and delivers training
Employee applies it
Evidence is sent back to L&D
L&D chases compliance
My model
L&D designs the behavior
Leader applies it in the work
Business leader inspects the operating metric
Behavior becomes part of management
The business verifies the behavior, keeping accountability with the operating line.
The participant's leader already held a regular business review. Clarity measures and visible commitment closure became part of that conversation, keeping accountability with the business line.
2. AI-enabled roleplay and feedback
Facilitator-bound practice
One facilitator's calendar
Limited repetitions per leader
Practice competes with delivery time
AI-enabled practice
Many leaders practicing in parallel
Repeated attempts, immediate feedback
Human calibration where it is actually needed
Practice capacity did the work here. The technology just made repetition cheap.
Before leaders used new scorecards, priority-reset language, or tradeoff conversations with their teams, they could rehearse against AI personas and digital voice tools tuned to the business context. The practice was low stakes, repeatable, and available on demand.
3. Peer-to-peer pods
Central model
Leader hits a barrier
Request routes to L&D
L&D schedules, reviews, responds
Distributed model
Triad reviews each other's commitments
Peers challenge unclear priorities
Troubleshooting happens closer to the work
What this produces is lateral accountability across functions.
4. Workflow automation: make the desired behavior easier than avoiding it.
I embedded the behavior into tools leaders already used, with automated prompts to document a cross-functional tradeoff or close a commitment where the work was already happening.
- Standalone L&D task
- → embedded operating habit
- → automatic evidence trail
- → less enforcement required
Phase 5 · Measure
Did anything actually change?
Same instrument, same rater-panel composition, six months later. 15 of 15 responded.
Measure
Movement at six months
Psychological Safety, cohort average of 12
9.810.4+0.6
Meaningful Work, cohort average of 12
10.410.7+0.3
Dependability, cohort average of 12
10.110.9+0.8
Structure and Clarity, cohort average of 12
9.610.9+1.3
Impact, cohort average of 12
10.510.8+0.3
Core Values, cohort average of 12
10.811.0+0.2
Same 15 leaders, same instrument, same scoring method. Six dimensions, three behavioral items each, rated 1 to 4 for impact and summed to a 3 to 12 dimension score; self ratings are excluded from the cohort figures. Approximately 90 percent of raters were unchanged, and replacement raters were matched by relationship type. The figures report observed movement on the repeated instrument across the matched cohort.
The operating evidence
- Peers stopped maintaining shadow trackers for Field Operations commitments
- Two peer leads named the monthly recap as the reason they stopped duplicating tracking
- Clarity practices survived a regional structure change in month five
- Closure rate is reported in the participant's own monthly business review
- Coaching support ended; the next contact was the reassessment
One direct report still rates clarity at 2, citing decision rights on vendor issues specifically. That is carried forward rather than averaged away.
Finishing the program was never the success condition. The behavior had to survive after the coach stepped away.
Measurement
What changed
9.6 → 10.9
Structure and Clarity, cohort average of 12, same instrument at six months
15
Leaders assessed at both points, approximately 90% of raters unchanged
6
Dimensions reassessed, three behavioral items each
Evidence
Artifacts for this case
Supplied samples are public-safe and sanitized.
360° Leadership Assessment (sanitized excerpt)
The six dimensions, the 1–4 behavioral scoring method, the perception-gap logic, and how the report feeds the debrief.
Format: PDF · fictional participant, sanitized data
Recreated sampleView sanitized sampleCompleted 90-day development plan
Baseline, two priorities, six observable commitments, 30/60/90 checkpoints, sustainment plan, and six-month reassessment.
Format: PDF · fictional participant, sanitized data
Recreated sampleView completed sampleSenior-leader facilitation excerpt
Session purpose, the five conditions as diagnostic lenses, a scenario with facilitator notes, and the 30-day commitment card.
Format: PDF · fictional scenarios and participant details
Recreated sampleView facilitation sample