Danielle Beram

Case study · Sandbox VR

Leadership development as an operating system

From insight to operating behavior.

Problem
Leadership feedback produced insight, then stalled because the operating rhythm had no mechanism to carry it forward.
Intervention
One closed loop: 360 diagnosis, senior-leader practice, observable 90-day commitments, verification inside the business review, reassessment at six months.
Result
Across the same 15-leader cohort, average Structure and Clarity ratings moved from 9.6 of 12 at baseline to 10.9 of 12 at six months, on the same instrument. Coaching support ended and the practices held.

Overview

What this was

The development cycle translated stakeholder perception into observable behavior, embedded that behavior into normal management routines, and returned six months later to test whether the experience of working with the leader had changed.

  • Designed the leadership framework and the 360 instrument built against it
  • Facilitated the senior-leader team effectiveness sessions
  • Ran every individual debrief and built the 90-day plans with participants
  • Moved verification into the participant's own business review
  • Built and administered the six-month reassessment

The system

Diagnosis, build, and how it ran

The constraint

The leaders already had plenty of feedback. What they lacked was a mechanism that carried insight into the quarter. Peers kept building shadow trackers. Direct reports kept re-deriving their own priorities each week.

So I built the cycle as one system with a single test at the end: does the experience of working with this leader change, and does the change survive after the coach steps away?

Phase 1 · Diagnose

The gap between intent and experience

A 360 usually arrives as a score. I use it to find where a leader's intent and their stakeholders' experience diverge, which patterns repeat across groups, and which few behaviors cause the most friction.

What the diagnosis surfaced

I read the report for divergence: where a leader's own rating and their stakeholders' ratings separate, which groups see it, and which few behaviors sit underneath the gap. The gap identifies where intent and experience differ.

What made it usable

  • Three lowest items named: one on expectations, one on priority stability, one on visible closure
  • One strength selected deliberately to help carry the change
  • One real gap, psychological safety, explicitly deferred: three priorities at once produce three partial results

Define

Vague development goal → observable operating behavior

This is the pivot the whole system depends on. Every commitment is written so a third party can tell whether it happened without asking the participant.

  • Written so it can be observed

    Issue any priority change in writing within 24 hours and name what is being dropped to make room.

  • Written so it can be observed

    Write a one-page role scorecard for each direct report (outcomes owned, decision rights held, criteria their work is judged against) and review it live with each person.

  • Written so it can be observed

    Post closure in the same thread where the commitment was made, on or before the date; if the date will slip, post a revised date before the original passes.

  • Written so it can be observed

    Send peer function leads a monthly written recap: what closed, what slipped, what changed.

Phase 2 · Practice

Facilitating senior leaders against evidence

The team effectiveness session for 20 executives, VPs, directors, and functional heads tests leadership identity claims against observable behavior and ends in a commitment written for third-party verification.

I use the five team conditions as diagnostic lenses, asking leaders to test whether their teams actually experience the conditions they believe they create.

Practice

Five conditions, five diagnostic questions

Select a condition to see the question I put to the room.

People can raise risk, admit mistakes, and challenge decisions without managing the leader's reaction.

When was the last time someone publicly challenged one of your decisions?

Holding the tension

When a senior leader says “people here can speak up, we're already an open culture,” I treat it as a hypothesis to test against the team's experience. Then I ask the room the question above and stop talking for eight to ten seconds.

If the answer is hard to retrieve, that is data. The session closes with each leader reading their committed behavior out loud to a partner. The partner tests one thing: could they recognize the behavior when it happens? Anything too vague gets rewritten.

Facilitation excerpt uses fictional scenarios and participant details. The original program deck remains Company Confidential.

Danielle Beram facilitating a leadership development session from a stage, with a slide titled Effective Leadership behind her and seated participants in the foreground.

What the session produced

The Project Aristotle senior-leader workshop is measured through documented follow-through after the session.

  • 20 participants
  • 19 documented a behavioral commitment
  • 16 completed the 30-day follow-up
  • 14 could point to a concrete application of the behavior in their work

Counts come from the session record and 30-day participant follow-up. They reflect self-reported follow-through.

Phase 3 · Turn insight into behavior

Feedback → behavior → evidence

Two priorities, six commitments, each with a cadence, an evidence source, and a measure. The evidence is captured where the work already happens.

Define

How a commitment is constructed

Select a step to see what it produces.

The checkpoints did real work

At 30 days the plan was at risk. Several role scorecards were still unwritten, and the missing ones belonged to the regions in the middle of a structure change, deferred on the reasoning that the roles might move.

The adjustment was to write them against the current role, dated, marked provisional. Clarity work gets postponed precisely when conditions are unclear, which is when it matters most. By 60 days the practices were holding without prompting and support moved from weekly to biweekly.

Phase 4 · Scale

Designed to run inside the business

Four mechanisms move reinforcement out of my calendar and into the business.

1. Manager-led verification

Traditional

  1. L&D designs and delivers training

  2. Employee applies it

  3. Evidence is sent back to L&D

  4. L&D chases compliance

My model

  1. L&D designs the behavior

  2. Leader applies it in the work

  3. Business leader inspects the operating metric

  4. Behavior becomes part of management

The business verifies the behavior, keeping accountability with the operating line.

The participant's leader already held a regular business review. Clarity measures and visible commitment closure became part of that conversation, keeping accountability with the business line.

2. AI-enabled roleplay and feedback

Facilitator-bound practice

  1. One facilitator's calendar

  2. Limited repetitions per leader

  3. Practice competes with delivery time

AI-enabled practice

  1. Many leaders practicing in parallel

  2. Repeated attempts, immediate feedback

  3. Human calibration where it is actually needed

Practice capacity did the work here. The technology just made repetition cheap.

Before leaders used new scorecards, priority-reset language, or tradeoff conversations with their teams, they could rehearse against AI personas and digital voice tools tuned to the business context. The practice was low stakes, repeatable, and available on demand.

3. Peer-to-peer pods

Central model

  1. Leader hits a barrier

  2. Request routes to L&D

  3. L&D schedules, reviews, responds

Distributed model

  1. Triad reviews each other's commitments

  2. Peers challenge unclear priorities

  3. Troubleshooting happens closer to the work

What this produces is lateral accountability across functions.

4. Workflow automation: make the desired behavior easier than avoiding it.

I embedded the behavior into tools leaders already used, with automated prompts to document a cross-functional tradeoff or close a commitment where the work was already happening.

  • Standalone L&D task
  • → embedded operating habit
  • → automatic evidence trail
  • → less enforcement required

Phase 5 · Measure

Did anything actually change?

Same instrument, same rater-panel composition, six months later. 15 of 15 responded.

Measure

Movement at six months

  • Psychological Safety, cohort average of 12

    9.810.4+0.6

  • Meaningful Work, cohort average of 12

    10.410.7+0.3

  • Dependability, cohort average of 12

    10.110.9+0.8

  • Structure and Clarity, cohort average of 12

    9.610.9+1.3

  • Impact, cohort average of 12

    10.510.8+0.3

  • Core Values, cohort average of 12

    10.811.0+0.2

Same 15 leaders, same instrument, same scoring method. Six dimensions, three behavioral items each, rated 1 to 4 for impact and summed to a 3 to 12 dimension score; self ratings are excluded from the cohort figures. Approximately 90 percent of raters were unchanged, and replacement raters were matched by relationship type. The figures report observed movement on the repeated instrument across the matched cohort.

The operating evidence

  • Peers stopped maintaining shadow trackers for Field Operations commitments
  • Two peer leads named the monthly recap as the reason they stopped duplicating tracking
  • Clarity practices survived a regional structure change in month five
  • Closure rate is reported in the participant's own monthly business review
  • Coaching support ended; the next contact was the reassessment

One direct report still rates clarity at 2, citing decision rights on vendor issues specifically. That is carried forward rather than averaged away.

Finishing the program was never the success condition. The behavior had to survive after the coach stepped away.

Measurement

What changed

  • 9.6 → 10.9

    Structure and Clarity, cohort average of 12, same instrument at six months

  • 15

    Leaders assessed at both points, approximately 90% of raters unchanged

  • 6

    Dimensions reassessed, three behavioral items each

Evidence

Artifacts for this case

Supplied samples are public-safe and sanitized.

  • 360° Leadership Assessment (sanitized excerpt)

    The six dimensions, the 1–4 behavioral scoring method, the perception-gap logic, and how the report feeds the debrief.

    Format: PDF · fictional participant, sanitized data

    Recreated sampleView sanitized sample
  • Completed 90-day development plan

    Baseline, two priorities, six observable commitments, 30/60/90 checkpoints, sustainment plan, and six-month reassessment.

    Format: PDF · fictional participant, sanitized data

    Recreated sampleView completed sample
  • Senior-leader facilitation excerpt

    Session purpose, the five conditions as diagnostic lenses, a scenario with facilitator notes, and the 30-day commitment card.

    Format: PDF · fictional scenarios and participant details

    Recreated sampleView facilitation sample