Engaging the Workforce to Deliver the Programme: Closing the Enactment Gap in Major Projects

Infrastructure workforce collectively rehearsing an approved plan before coordinated delivery on site

Table of Contents

Last Updated on August 23, 2026

Dr Ben Guy, Urban CGI — Urban Consulting Group International

Bent Flyvbjerg explained how big things get done. This article is about how they get built — by the tens of thousands of people who have to enact a plan somebody else approved.

That distance, between an approved plan and the coordinated behaviour of the workforce and supply chain who must perform it, is where the money goes. The industry does not name it, cannot price it, and has never built an instrument for it. I call it the enactment gap, and engaging the workforce to close it is the last real lever a project director holds once the design is settled and the consents are in.

This is written for the people carrying a capital programme and a baseline they are accountable for holding. Everything below is about where the money actually leaks and what can be demanded to stop it.

Infrastructure workforce collectively rehearsing an approved plan before coordinated delivery on site

The most expensive sentence in British infrastructure

In July 2025, giving evidence to the House of Commons Transport Committee, the chief executive of HS2 said something that should have changed the conversation and did not.

Mark Wild had run Crossrail through its recovery and had been at HS2 for seven months. Asked why civil engineering works budgeted at £19.5bn had consumed £26bn while barely passing the halfway mark, he did not talk about the estimate:

“The projects must not be mobilised and commenced if you haven’t got the design and consents, because the productivity of the teams, the hard working teams, is so leveraged if you’re waiting for the design.”

And then, more plainly: “I think when all said and done, the cost exceedence though has mostly been inefficiency of work, because we started too soon.”

Note where he locates the loss. Not in the business case. Not in the forecast. In the hardworking teams — willing people, working inefficiently, on a project that was not ready for them. In the same session he named the mechanism: “If you lose control of the programme, you end up at the extreme end of optimism bias, which ends up in delusion.”

Wild is describing something the industry has no measurement for and no method against. It needs a name, so I will give it one: the enactment gap.

What Flyvbjerg proved, and it is a Big Deal

Any serious argument about why major projects fail starts with Bent Flyvbjerg, because he built the dataset. It holds more than 16,000 projects across more than twenty project types in 136 countries, and from it comes the most quoted finding in the field: 8.5% of projects hit both cost and time, and 0.5% hit cost, time and benefits together. He named the pattern the iron law of megaprojects — over budget, over time, under benefits, over and over again.

The granular figures are harder still. Rail runs a mean cost overrun of 44.7% alongside a mean demand shortfall of 51.4%, and costs are underestimated in almost nine out of ten projects. And this, which should end every conversation about industry learning: “cost underestimation has not decreased over time. Underestimation today is in the same order of magnitude as it was 10, 30, and 70 years ago.”

His diagnosis is behavioural and it separates two root causes. Delusion is optimism bias — honest underestimation by people who know perfectly well that comparable projects ran late. Deception is strategic misrepresentation, and in 2002 he put it without hedging: underestimation “cannot be explained by error and is best explained by strategic misrepresentation, that is, lying.” Against the first he prescribes reference class forecasting, which the UK Department for Transport adopted in 2004 and still mandates. Against the second, governance.

This is more empirical rigour than the rest of the field combined. It is also, by his own explicit design, a theory that stops at the site gate.

Where the theory stops: nobody has studied enactment

Flyvbjerg has answered this in print, consistently: “In behavioral terms, the causal chain starts with human bias which leads to underestimation of scope during planning which leads to unaccounted for scope changes during delivery which lead to cost overrun.” Or in the book’s compressed form: projects don’t go wrong, they start wrong.

On that account, delivery is where failure becomes visible, not where it is made. Everything downstream of the decision to proceed — the sequence, the crew, the supply chain, the supervisor, the fifty-thousand-page method statement — is a transmission mechanism for an error committed upstream. It follows that you fix projects in the planning room.

His corpus reflects that with integrity, and it contains no empirical study of workforce behaviour, site-level culture, supervision, competence-building, or how a tier-three subcontractor comes to perform a sequence correctly at four in the morning. There could not be one: the unit of analysis is the project, and the database holds no workforce-level variable. There is a single chapter on the delivery organisation, arguing for integrated teams and psychological safety, and its exemplar is Heathrow Terminal 5 — research done by other people in a different tradition.

The field’s own systematic review says it more bluntly. Denicol, Davies and Krystallis screened 6,007 titles and read 86 full papers, and concluded: “What is missing in current research is an understanding of megaprojects as a complete production system — from planning, through design, manufacturing, and construction, to integration and handover to operations.”

Nobody has studied enactment. That is what this article is about.

The enactment gap, defined

The enactment gap is the distance between an approved plan and the coordinated behaviour of the people who must perform it. Not a scheduling problem, a documentation problem or a training problem, though it presents as all three. It is the difference between a plan that exists and a plan that happens.

What the enactment gap costs, and who has already itemised it

Read the audit reports for the mechanism rather than the headline and the same failure appears in four currencies.

Crossrail. Of £2.5bn of cost increase between 2013 and 2018, design changes accounted for £714m. Compensation event notices reached 21,000 by January 2016 with 1,800 awaiting assessment, at which levels the National Audit Office observed that “time [is] absorbed managing commercial elements of the contract, rather than carrying out productive work.” The handover schedule was “an aspirational plan designed to improve progress by suppliers, rather than to provide a reality check”, with contractors’ planned productivity bearing “little resemblance to the historic progress that contractors had made”. Even after the reset, contractors met only 30% of milestones through 2019–20 while management “continued to base its plans on more optimistic levels of productivity.”

The Big Dig. Change orders were finalised at 14% to 19% over original contract prices against a contingency budgeted at 7%. Across thirteen contracts reviewed by the Massachusetts Inspector General, more than 1,600 change orders took $310m of original value to $435m. Flyvbjerg scores the project at a 220% real cost overrun.

Berlin Brandenburg. Terminal capacity was doubled during construction. Thirty to forty companies were engaged rather than a single general contractor, so nobody owned integration. The airport opened nine years late at €7.3bn against €2.83bn, with 550,000 identified defects.

Sydney CBD and South East Light Rail. A dispute over electricity cable information on George Street produced an A$576m settlement. Final cost A$3.147bn against an original business case of A$1.6bn, a 97% overrun, and the Auditor-General found expected benefits had not been updated since 2015.

None of those four is a forecasting error. Every one of them is something that happened to the work after the money was committed: a change order, a claim, an interface nobody owned, a piece of information about a cable that did not reach the people digging.

And then the statistic that should be the most discussed in the built environment and is barely known within it. Goolsbee and Syverson, on seventy years of US Census and Bureau of Labor Statistics data, find that a construction worker in 2020 produced less than a construction worker in 1970. Between 1950 and 2020 manufacturing labour productivity rose ninefold; construction labour productivity fell below its 1950 level, on capital investment that rose nearly eightfold. Of twenty-nine OECD countries studied, sixteen had negative average construction labour productivity growth between 1996 and 2019.

Every comparable industry took the productivity dividend of the twentieth century. Construction did not. Consider what construction has that manufacturing does not: the work is performed by a different, temporary, multi-employer crowd every time, in a place that has never existed before, from instructions written down.

The enactment gap in other sectors

The pattern is not a rail pattern, and it is not a British one. Wherever an auditor has looked closely at a failed programme, the same three findings appear: the design was not settled, the supply chain was not capable, and nobody had engaged the workforce in what they were about to be asked to do.

Nuclear: incomplete design, immature supply chain, untrained workforce

The US Department of Energy has written the enactment gap into its own policy document. On Plant Vogtle — US$14.3bn certified in 2008–09, roughly US$35bn delivered, about seven years late — it says: “Vogtle began construction with an incomplete design, an immature supply chain, and an untrained work force.”

Three failure modes, one sentence, about the flagship project of a national industrial strategy. The detail underneath is worse. The National Academies record a stop-work order at the Lake Charles module factory in 2010 after inspectors found “inferior welds and welders accepting welds that they had not made”, after which many factory welds “had to be completely redone”.

Finland’s regulator saw the same thing at Olkiluoto 3 in 2006, years before most of the overrun had happened. From STUK’s investigation report: the construction plan was incomplete, so “the manufacturer has not always known what the most recent updated version of the design drawings is.” The concrete supplier “had no experience in a nuclear power plant construction prior to the OL3 project”, and — this is the sentence to sit with — “no training was provided to the staff involved in the fabrication of concrete concerning practices in the nuclear field and the safety significance of their work.” Olkiluoto came in at roughly €11bn against a €3bn turnkey contract, thirteen years late.

France reached the same diagnosis at Flamanville 3, now recalculated by the Cour des comptes at €23.7bn against an original €3.3bn. The Folz report named immature design studies at launch, insufficient local supervision, and a generalised loss of skills after twenty years without building a reactor. And Hinkley Point C’s most recent slip was attributed by EDF not to design and not to finance, but to “lower than expected productivity on the vast electromechanical installation programme” — the rate at which people install things.

Set that against the technologies that do not overrun. Across 11,011 projects in 126 countries, mean cost outturn as a ratio of budget runs 1.01 for solar power, 1.08 for transmission and 1.12 for wind, against 2.20 for nuclear. Read the medians and the mechanism is exact: solar, transmission and wind have a typical outcome of on budget. A solar farm is one panel repeated two hundred thousand times, so the design is finished before mobilisation and the crew learns the sequence on the first hundred. Modularity is usually explained as a procurement idea. It is better understood as an engagement idea — repetition is what lets a plan be transmitted into a workforce and then improved by it.

Mining and resources: assurance is not readiness

Snowy 2.0 was announced at A$2bn in 2017, contracted at A$5.1bn in 2019 and reset to A$12bn in 2023, with completion moved to December 2028. Construction proceeded with the design immature. In February 2023 the tunnel boring machine Florence struck soft ground and stayed stuck for about eleven months.

Two findings from the audits deserve wider attention than they have had. The Australian National Audit Office reported in 2026 that management since the reset was “partly effective” with “significant deficiencies”, and that “the baseline schedule for the project’s completion had not been agreed upon” — so the auditors could not determine how far from completion it was. And the earlier audit, tabled in June 2022, had found governance of early implementation “effective”, supported by a geotechnical baseline built from deep boreholes and testing. Eight months later Florence hit ground that baseline had not anticipated.

Governance assurance is not delivery readiness. An auditor auditing process will pass a project that is about to fail in the tunnel.

The same shape appears across Australian LNG. On figures compiled from company disclosures, every project in the growth wave exceeded its capital guidance and every project started production later than its schedule guidance — Pluto 33%, Gorgon 26%, Gladstone 32%, Ichthys 82%. And the stated causes are all things that happened to the work rather than errors in the forecast: Pluto replacing the flare and insulation during commissioning, Wheatstone’s late modules and underestimated quantities, Ichthys’s late vessel and scope disputes.

Housing: the same gap at national scale

England recorded 208,600 net additional dwellings in 2024–25, a 6% fall on the year before, against a government-assessed national need of 370,408. Australia completed roughly 232,000 net new dwellings in the first eighteen months of an Accord targeting 1.2 million in five years, with the official supply council now projecting around 980,000.

Where an independent body has given a diagnosis, it points at the same place. The Competition and Markets Authority went looking for a market-power explanation in 2024, explicitly rejected land banking as a root cause, and concluded that “the planning system is exerting a significant downward pressure on the overall number of planning permissions being granted”. Australia’s National Housing Supply and Affordability Council puts it in one line: “the administration of planning and building controls for new housing is unnecessarily complex, causing significant delays.”

Honesty requires the counter-note. The cyclical swing in housing starts is demand-side — rates, tax and finance. Delivery capacity sets the ceiling; demand decides where you sit under it.

Why documents and RAMS cannot engage a workforce

The Lower Thames Crossing development consent application cost £267m to prepare and ran to 359,866 pages across 2,383 documents: 94,534,273 words. At two hundred words a minute, around the clock without stopping, it would take 328 days to read. That is about £111m of planning cost per mile of tunnel — before a metre of it was built.

The trend is the same everywhere the consenting burden is measured. Time to a decision on a Nationally Significant Infrastructure Project rose from 2.6 years in 2012 to 4.2 years in 2021, a 62% increase in nine years, on the government’s own figures. Around 22% of development consent decisions have been legally challenged, and National Highways calculates the additional cost of legal challenge at between £66m and £121m per scheme. HS2 required more than 7,000 individual consents for the civil engineering alone.

We have industrialised the production of plans and left the transmission of them exactly where it was. A method statement is a document addressed to a reader who will not read it, describing a physical performance in a medium that cannot show a hand position, a sightline, a standing place, or the half-second before a lift goes wrong. It is written to survive an audit, which is a different design brief from being understood by the person who has to do the work.

The paperwork is not the plan. It is the record of the plan. The plan is what people do.

Plans and documents becoming shared workforce rehearsal before a coordinated team enters an infrastructure workface

Engaging the workforce through more modalities

Here the behavioural science is better than its reputation, and more specific than the popular version.

Demonstration transmits the shape of a movement far more powerfully than it improves the result of one. Ashford, Bennett and Davids, meta-analysing observational modelling across task types, found an effect size of d = 0.77 on movement dynamics — form and coordination — against d = 0.17 on movement outcome. The effect was largest of all for serial tasks: multi-step sequences, at d = 1.62. A construction method statement is a serial task. So is a lift, a track possession, a confined-space entry and a shift handover.

The advantage of showing over telling is modest in general and large in exactly this case. Höffler and Leutner, across twenty-six studies and seventy-six comparisons, found dynamic representation beat static at d = 0.37 overall, and at d = 1.06 for procedural-motor knowledge.

Modelled behaviour also lasts differently from modelled information. Taylor, Russ-Eft and Chan’s meta-analysis of 117 studies of behaviour modelling training found that declarative knowledge decayed over time while effects on skills and job behaviour “remained stable or even increased”. Three of their transfer findings deserve to be read twice by anyone running a large workforce: transfer was greatest when negative as well as positive models were shown, when practice used scenarios the trainees themselves generated, and when the trainees’ superiors were trained too.

That last point is the cascade, and it matches the strongest correlate of harm in the organisational literature. Christian and colleagues found that group safety climate — the shared, workgroup-level reading of what the supervisor actually rewards and tolerates — had the strongest association with accidents and injuries of any variable tested, ahead of individual attitude, knowledge and motivation.

One necessary correction, against the grain of my own industry’s marketing. Passive watching is not the intervention. Burke and colleagues, across ninety-five studies and 20,991 workers, graded safety training by engagement and found knowledge gains rising from d = 0.55 for the least engaging methods to d = 1.46 for the most. Lectures, pamphlets and videos sit in the least engaging category. Behavioural modelling with substantial practice and dialogue sits in the most. Their conclusion cuts against the drift of the entire e-learning sector: the findings “challenge the current emphasis on more passive computer-based and distance training methods.”

So: demonstration transmits sequence. Modelling transmits behaviour, and it cascades from supervisors. Neither works passively — the observer has to do something with what they saw. On mechanism, the popular account reaches for mirror neurons, and it should reach more carefully: matching neurons were recorded in macaque premotor cortex from 1992, but direct single-neuron evidence in humans remains essentially one study, in regions that are not the homologues of the macaque areas, and the interpretive claims have been contested for fifteen years. What survives the criticism is a contribution to imitation. The behavioural findings above do not depend on the neuroscience being settled, and I would rather build on the meta-analyses.

Terminal 5: what poor staff engagement cost a project delivered on budget

Heathrow Terminal 5 was delivered for £4.3bn, on time and to the target cost it was sanctioned against in 2002 — which, given the base rates above, is a genuine achievement and worth understanding.

How it was done reads like a list of ways to settle the work before mobilising. BAA held the major risks itself and insured the programme, deliberately to strip risk pricing out of tender bids. Normally competing contractors were co-located from the start in integrated teams, with a dedicated organisational effectiveness team brokering relationships between firms. Every worker on the project received a @t5.co.uk email address regardless of who actually employed them; as the Manchester case study records it, “there weren’t lots of people on site, there were T5 people on site.” The programme ran “first run studies” to validate production methods before committing to them, moved work off site into pre-assembly, and ran two funded logistics consolidation centres feeding the site just in time.

Then it opened. On 27 March 2008, 36,584 passengers were affected, 23,205 bags required manual sorting, and over the following ten days more than five hundred flights were cancelled and roughly 42,000 bags failed to travel with their owners. The Transport Committee’s verdict: “what should have been an occasion of national pride was in fact an occasion of national embarrassment.”

Its diagnosis is the reason this case matters more than any other in this article:

“Most of these problems were caused by one of two main factors: insufficient communication between owner and operator, and poor staff training and system testing.”

And on the trials, which were not trivial — 66 of them, 15,000 volunteers, 400,000 bags:

“The proving trials may have succeeded in identifying improvements … but they failed in the ultimate objective of getting the system to a point where it worked well enough to cope with the opening successfully.”

British Airways staff described three days of familiarisation across an area the Committee’s witnesses compared to the size of Hyde Park, with hands-on practice unavailable because the building was still a construction site. And BAA’s own chief executive told the Committee what had happened to the integration that built the asset: “during construction of Terminal 5 that appeared to be the case. Around about or just prior to the opening … that togetherness deteriorated.”

There it is. The most sophisticated delivery organisation in the history of British construction rehearsed the building and did not rehearse the operation. The @t5.co.uk identity covered the people who designed and built it. It did not cover the people who would run it. A £4.3bn asset delivered to target was compromised by a readiness failure costing tens of millions and an enormous amount of reputation — and the marginal spend that would have prevented it was trivial against the capital cost. Andrew Davies, one of the authors of the celebrated 2009 account of T5’s systems integration, returned to the case in 2010 under the title “From hero to hubris”. That is an unusually honest piece of scholarship, and the honesty is the finding.

What the wins have in common: settled work and engaged teams

The argument is falsifiable, so test it against success.

Madrid built 131 km of metro and 76 stations in eight years from 1995, at around US$16.6m per kilometre for tunnelling, with major projects averaging thirty-two months each. Manuel Melis Maynar, who ran it, did not attribute that to forecasting. He attributed it to a client organisation of nine regional engineers and four geotechnical specialists making decisions in-house, on the grounds that with external project managers “you lose months”. He weighted tenders at 50% technical content and staffing, 30% cost and 20% programme — price under a third of the score. He refused untested technology. And he ran standard, repeated station designs so that each team learned from the last. His own formulation is worth keeping: “a well-executed tunnelling project is an art; the client should be prepared to spend the necessary time in choosing the artist.”

Bilbao is the other proof, and it is the one closest to our own work. The Guggenheim came in at US$97m against a US$100m budget — the most formally complex building of its generation, delivered by an architect with no track record of large-budget discipline. Flyvbjerg’s description of how: Gehry’s team “spent two years thinking through and simulating every detail, in effect building the museum on computers before they built it in reality.”

And then the inversion. Crossrail overran by billions and by up to three and a half years, and the railway it produced carries around 800,000 journeys a day, holds the highest customer satisfaction of any Transport for London line at 83%, and won the RIBA Stirling Prize in 2024. It is the mirror image of Terminal 5. T5 got the asset right and the engagement wrong; Crossrail got the delivery wrong and the operational readiness right. Two projects, one country, one sector, opposite failure modes — and in both cases the variable that decided how it felt to the public was whether the people were ready, not whether the forecast was accurate.

Pixar planning, CGI planning and engagement

Chapter 4 of How Big Things Get Done is titled “Pixar Planning”, and the prescription in it is unambiguous:

“When a minimum viable product approach isn’t possible, try a ‘maximum virtual product’ — a hyperrealistic, exquisitely detailed model like those that Frank Gehry made for the Guggenheim Bilbao and all his buildings since and those that Pixar makes for each of its feature films before shooting.”

And on Gehry: “only after the completion of its ‘digital twin’ … did construction begin in the real world.”

He is right, and the industry has half-heard him. The maximum virtual product the sector actually built is a model of the asset. BIM, digital twins, clash detection, federated models: an exquisite representation of the thing, to a standard that would have astonished anyone working in 1990.

But the maximum virtual product for a project is not a model of the asset. It is a model of the work. The sequence, the exclusion zone, the lift, the traffic switch, the hold point, the shift handover, the person walking backwards guiding a reverse. Pixar does not model the set. It models the performance — as many as eight full iterations of a film before a frame is animated, because iteration in development is cheap and iteration in production is not.

Pixar planning names the principle. What we call CGI planning is what the principle looks like when the product is a construction sequence rather than a film: the method rehearsed in the real geometry, in real time, with the people who will perform it, before anyone is standing in it. Urban CGI planning is the practice of building the performance before you build the asset. The distinction is not technological. Both use the same tools. One models what will exist; the other models what will happen.

Almost nobody in infrastructure models the performance, which is a strange gap to have left open, because the industry accepts the logic everywhere else. Aviation rehearses. Surgery rehearses. The military rehearses. Film rehearses and calls it pre-visualisation. Construction issues a document and holds a briefing.

What strong engagement looks like, measured

The published evidence above is about the size and shape of the gap. The question a director will reasonably ask next is whether anything closes it, and the honest answer is that the industry has almost no instrumented cases, because almost nobody has tried at scale and fewer still have measured.

Here is one of ours, and I state it as first-party evidence rather than published research. On a large minerals processing operation, the operator’s own work-from-heights record before we were engaged ran to roughly one fatality every three years, costed internally at around $25m each. We simulated and rehearsed the work-from-heights procedures for what became the longest shutdown in that plant’s history — in particular the selection and use of fall restraint over fall arrest. Zero work-from-heights incidents across the shutdown.

That distinction is worth dwelling on, because it is the whole argument in miniature. Restraint stops a person reaching the edge. Arrest catches them after they have gone over it. The difference is one line in a procedure, universally understood in the abstract, and routinely got wrong in the field — because getting it right is a matter of anchor points, lanyard lengths, standing positions and where a person’s feet actually are on a particular structure at a particular moment. That is not information. It is a performance, and it is precisely the class of problem demonstration solves and documents do not.

Take the result for what it is. One operation, uncontrolled, no counterfactual, and the confounds in any shutdown are numerous. It is not a study. But it is a number, it is ours, and it is the kind of number the industry does not collect: an operator with a known baseline, a longer and more complex delivery event than it had ever run, and an incident record that moved in the opposite direction to the exposure.

The reason it is worth stating even in that unsatisfying form is that the alternative evidence base is worse. There is no serious research literature on operational readiness at all — the discipline exists in practitioner manuals and airport commissioning handbooks, and almost nowhere in the peer-reviewed record. The absence is itself the finding.

Four things a project director can do to engage the workforce

I will state the conditions rather than the method, because the method is our work. But these are not proprietary, and any client can write them into a scope, an assurance regime or a tender.

One: demand that it is seen, not read. The transmission medium has to match what is being transmitted. A serial physical sequence transmits by demonstration at roughly four times the effect size at which it transmits by outcome description, and no amount of document quality changes that.

Two: demand the actual place. A generic animation teaches a generic sequence, and no site is generic. The failures above are overwhelmingly failures of interface and context — utilities under George Street, three signalling systems and half a million assets on Crossrail, nineteenth-century services under Boston, soft ground fifteen kilometres into a tunnel in the Snowy Mountains. If the rehearsal is not in the real geometry with the real constraints, it rehearses a project that does not exist.

Three: demand that it is tested under load before anyone is standing in it. Not the nominal case. The delivery that arrives early, the crowd that goes the wrong way, the weather, the plant that is where it should not be. This is where agent-based simulation earns its keep, and where the honest limit sits too: simulation effect sizes attenuate as you move from the training environment toward the real world, from around 1.1 in the simulator to 0.50 on real-world outcomes. That attenuation is a reason to measure it, not a reason to skip it.

Four: demand that enactment is evidenced, not assumed. The question to put to a delivery organisation is not whether the procedure was issued and the briefing held. It is who has demonstrated they can perform it. Almost no major project can answer that today at the level of the individual, which is remarkable given what the industry can answer about the location of a bolt. Terminal 5 could tell you the position of every one of its components. It could not tell you which baggage handlers had ever run the system with real luggage in it.

Those four are auditable. That is the point of stating them. A director who writes them into an assurance regime has changed what the delivery organisation is obliged to prove, which is the only lever that has ever moved this industry.

Why engagement matters more in the 2020s and 2030s than it did in the 2010s

Three facts about the next fifteen years.

The pipeline is unprecedented. McKinsey puts cumulative required infrastructure investment for 2025 to 2040 at US$106 trillion, with around two-thirds of it in Asia.

The workforce does not exist. Infrastructure Australia’s market capacity work puts the current infrastructure workforce at 204,000 against peak demand of 521,000: a shortage of 141,000, rising to around 300,000 by mid-2027, with regional shortages expected to quadruple. In the United States, 41% of the construction workforce is expected to retire by 2031 and only 10% is under twenty-five.

And technology is not coming to the rescue on site this decade. On-site construction robotics represented less than 0.03% of global construction spend in 2025. A survey of 951 contractors found artificial intelligence running at 45% for office and administrative functions, 23% for estimating and 20% for design and preconstruction. AI in construction today is a back-office technology.

Put those together and the binding constraint on the largest building programme in human history is the capability of human beings, arriving faster than any existing method can make them competent, on projects more complex than the ones we already fail at. That is an engagement problem, and it gets harder every year the pipeline grows.

Engaging the workforce is the delivery strategy

The language is already in circulation among the people who have recovered failing programmes. In his Smeaton Lecture at the Institution of Civil Engineers, Mark Wild named the binding constraint on delivery: “long-term planning is essential, allowing time to develop modular solutions, the workforce and appropriate technology.” Not the design, and not the money. The workforce, alongside the method and the technology, as something that has to be developed rather than assumed.

Notice what workforce engagement is not, in that framing. It is not a communications plan, a culture programme, or a set of values on a hoarding. It is a delivery strategy, aimed at the only variable still open once the design is settled and the consents are in — whether the people on site, internal and supply chain, can actually perform the plan.

Infrastructure teams coordinating interdependent tasks across a railway workface as one delivery sequence

Those people are the crew on nights, the banksman, the driver, the supervisor in their first week. The concrete workers at Olkiluoto who were never told why any of it mattered. The welders at Lake Charles signing off welds they had not made. The baggage handlers at Terminal 5 with three days on a building site and an area the size of Hyde Park to learn. They are not a transmission mechanism for a plan made somewhere else. They are where the plan either happens or does not.

Flyvbjerg is right that projects start wrong. They also finish wrong, for reasons that begin at the site gate and have never been measured with anything like the same rigour. Between the approved plan and the performed plan is a gap the industry does not name, cannot price, and has never built an instrument for.

Closing it is the interesting work of the next decade, and it is where we work. The wider argument behind it — that we plan for humans, the people who will use what we build as well as the people who build it — runs through our work on advanced pedestrian simulation.

Frequently asked questions

What is the enactment gap?

The enactment gap is the distance between an approved plan and the coordinated behaviour of the workforce and supply chain who must perform it. It is not a scheduling problem, a documentation problem or a training problem, although it presents as all three. It is the difference between a plan that exists and a plan that happens.

What is workforce engagement on a major project, and what is it not?

It is a delivery strategy aimed at the last variable still open once the design is settled and the consents are in: whether the people on site, internal and supply chain, can actually perform the plan. It is not a communications plan, a culture programme, a values exercise or a poster campaign. The test of engagement is not whether people have been told. It is whether they can do it, and whether anyone can show that they can.

Isn’t cost overrun already explained by optimism bias?

Partly, and Flyvbjerg’s evidence is unmatched: 8.5% of projects hit both cost and time, 0.5% hit cost, time and benefits together. But his causal model explicitly locates the root cause before delivery begins, and treats scope change and disruption during construction as downstream symptoms of upstream bias. That leaves the question of how a delivery organisation is made capable of performing the plan unanswered — and the field’s own systematic review says the missing research object is the megaproject as a complete production system.

Where does the money leak between the plan and the site?

Into change, claims and rework after the money is committed. Design changes accounted for £714m of Crossrail’s £2.5bn increase between 2013 and 2018, with 21,000 compensation event notices logged by January 2016. Big Dig change orders were finalised at 14% to 19% over original contract prices against a 7% contingency. Hinkley Point C’s most recent slip was attributed by EDF to “lower than expected productivity on the vast electromechanical installation programme”.

Why did Terminal 5 open badly if it was delivered on budget?

Because the asset was rehearsed and the workforce was not. T5 was delivered for £4.3bn against the target cost set at its 2002 sanction, using integrated teams, first-run studies and off-site pre-assembly. The Transport Committee found the opening failure was caused by “insufficient communication between owner and operator, and poor staff training and system testing”, and that 66 proving trials using 15,000 volunteers and 400,000 bags “failed in the ultimate objective”. The engagement that built the terminal stopped at the handover line.

Does a clean governance audit mean the workforce is ready?

No. The Australian National Audit Office concluded in June 2022 that governance of Snowy 2.0’s early implementation was “effective”, supported by a geotechnical baseline built from deep boreholes and testing. Eight months later the tunnel boring machine struck ground that baseline had not anticipated and stayed stuck for about eleven months. Governance assurance is not delivery readiness, and neither is evidence of engagement.

Can a method statement or a RAMS document engage a workforce?

Not well. Meta-analysis of observational modelling finds demonstration produces an effect size of d = 0.77 on movement form against d = 0.17 on movement outcome, and the largest effects of all — d = 1.62 — on serial, multi-step tasks. Every method statement describes a serial task. For scale, the Lower Thames Crossing consent application ran to 94,534,273 words, which at two hundred words a minute would take 328 days to read continuously. A document written to survive an audit is not the same artefact as one written to be enacted.

Is watching a video enough to engage a workforce?

No, and the evidence is blunt about it. Burke and colleagues, across 95 studies and 20,991 workers, found knowledge gains rising from d = 0.55 for the least engaging training methods to d = 1.46 for the most — and lectures, pamphlets and videos sit in the least engaging category. Behavioural modelling with substantial practice and dialogue sits in the most. Their conclusion cuts against the whole e-learning sector: the findings “challenge the current emphasis on more passive computer-based and distance training methods.”

What Disney and Pixar are genuinely expert at is holding attention and making a sequence legible, and that craft matters enormously in getting a method understood. But attention is the entry ticket, not the outcome. The observer has to do something with what they saw. On mechanism, the popular account reaches for mirror neurons and should reach more carefully: matching neurons were recorded in macaque premotor cortex from 1992, direct single-neuron evidence in humans remains essentially one study in regions that are not the homologues of the macaque areas, and the interpretive claims have been contested for fifteen years. What survives the criticism is a contribution to imitation. The behavioural findings above do not depend on the neuroscience being settled.

What is CGI planning, and how is it different from BIM or a digital twin?

BIM and digital twins model the asset: a representation of the thing that will exist. CGI planning models the work, and the people doing it — the sequence, the exclusion zone, the lift, the traffic switch, the hold point, the shift handover, the person walking backwards guiding a reverse. Both use the same tools. One models what will exist; the other models what will happen. Flyvbjerg calls the underlying principle Pixar planning, and describes Frank Gehry’s team on the Guggenheim Bilbao as having “spent two years thinking through and simulating every detail, in effect building the museum on computers before they built it in reality.” That building came in at US$97m against a US$100m budget.

How can a project director engage the workforce to deliver the programme?

Four things, and all four are auditable. Demand that the method is seen rather than read. Demand that the rehearsal happens in the real geometry with the real constraints, not in a generic animation. Demand that it is tested under load — the early delivery, the weather, the plant in the wrong place — before anyone is standing in it. And demand that engagement is evidenced rather than assumed: not whether the procedure was issued and the briefing held, but who has demonstrated they can perform it.

What does strong engagement look like, and can it be measured?

It looks like a workforce that can perform a sequence it has already seen, in the place it will happen, under conditions worse than the nominal case — and a record, at the level of the individual, of who has demonstrated that. Almost no major project can answer the last part today, which is remarkable given what the industry can answer about the location of a bolt. Terminal 5 could tell you the position of every one of its components. It could not tell you which baggage handlers had ever run the system with real luggage in it.

Does rehearsal transfer to real performance?

Partly, and the honest number is the attenuated one. In the healthcare literature, which is the best evidenced, simulation-based training shows effect sizes of around 1.1 measured in the simulator, around 0.8 on behaviour with real patients, and 0.50 on patient outcomes against no intervention. Against alternative well-designed instruction the advantage narrows to 0.30 to 0.66 for knowledge and skills. Effect attenuates as you move toward the real world. That is a reason to measure it, not a reason to skip it.

Bibliography

Megaproject performance and behavioural cause

Denicol, J., Davies, A. and Krystallis, I. (2020) “What Are the Causes and Cures of Poor Megaproject Performance? A Systematic Literature Review and Research Agenda”, Project Management Journal, 51(3), pp. 328–345.

Flyvbjerg, B. (2014) “What You Should Know About Megaprojects and Why: An Overview”, Project Management Journal, 45(2), pp. 6–19.

Flyvbjerg, B. (2021) “Make Megaprojects More Modular”, Harvard Business Review, November–December.

Flyvbjerg, B., Ansar, A., Budzier, A., et al. (2018) “Five things you should know about cost overrun”, Transportation Research Part A, 118, pp. 174–190.

Flyvbjerg, B., Budzier, A., Aaen, J., Keil, M. and Zottoli, M. (2026) “The Uniqueness of IT Cost Risk: A Cross-Group Comparison of 23 Project Types”, Project Management Journal, 57(1), pp. 14–43.

Flyvbjerg, B., Garbuio, M. and Lovallo, D. (2009) “Delusion and Deception in Large Infrastructure Projects”, California Management Review, 51(2), pp. 170–193.

Flyvbjerg, B. and Gardner, D. (2023) How Big Things Get Done. New York: Currency.

Flyvbjerg, B., Holm, M. S. and Buhl, S. (2002) “Underestimating Costs in Public Works Projects: Error or Lie?”, Journal of the American Planning Association, 68(3), pp. 279–295.

Delivery organisation, systems integration and readiness

Brady, T. and Davies, A. (2010) “From hero to hubris — Reconsidering the project management of Heathrow’s Terminal 5”, International Journal of Project Management, 28(2), pp. 151–157.

Davies, A., Gann, D. and Douglas, T. (2009) “Innovation in Megaprojects: Systems Integration at London Heathrow Terminal 5”, California Management Review, 51(2), pp. 101–125.

Gil, N. and Ward, D. (2011) Leadership in Megaprojects and Production Management: Lessons from the T5 Project. CID Technical Report No. 1, University of Manchester.

Melis Maynar, M., interviewed in “Model guides metro expansion”, Railway Gazette International, 1 January 2001.

Audits, inquiries and official reviews

Audit Office of New South Wales (2020) CBD South East Sydney Light Rail: follow-up performance audit, 11 June.

Australian National Audit Office (2022) Governance of the Snowy 2.0 Project, Auditor-General Report No. 33 of 2021–22, 15 June.

Australian National Audit Office (2026) Delivery of Snowy 2.0, Auditor-General Report No. 39 of 2025–26, June.

Competition and Markets Authority (2024) Housebuilding market study: final report, 26 February.

Cour des comptes (2025) Le programme d’EPR, 14 January.

Department for Transport (2025) TAG Unit A1.2: Scheme Costs, May.

Federal Highway Administration (2000) Review of Project Oversight and Costs: Federal Task Force on the Boston Central Artery/Tunnel Project, 31 March.

Folz, J.-M. (2019) La construction de l’EPR de Flamanville, report to EDF, October.

House of Commons Transport Committee (2008) The opening of Heathrow Terminal 5, HC 543, 3 November.

House of Commons Transport Committee (2025) Delivering major infrastructure: learning from HS2, oral evidence HC 1139, 9 July.

Massachusetts Office of the Inspector General (2003) A Big Dig Cost Recovery Referral, December.

National Academies of Sciences, Engineering, and Medicine (2023) Laying the Foundation for New and Advanced Nuclear Reactors in the United States.

National Audit Office (2019) Completing Crossrail, May; and (2021) Crossrail: a progress update, HC 299, 9 July.

National Housing Supply and Affordability Council (2026) State of the Housing System 2026, 20 April.

STUK (Finnish Radiation and Nuclear Safety Authority) (2006) Investigation Report 1/06, translation 1 September.

Stewart, J. (2025) Major Transport Projects Governance and Assurance Review: The HS2 Experience, Department for Transport, 18 June.

United States Department of Energy (2024) Pathways to Commercial Liftoff: Advanced Nuclear, September.

Productivity, error and delivery economics

Australasian Centre for Corporate Responsibility (2023) Australia’s LNG growth wave — did it wash for shareholders?, 27 November.

Britain Remade (2024) How the Lower Thames Crossing is breaking records for all the wrong reasons, 12 January.

Eash-Gates, P., Klemun, M., Kavlak, G., McNerney, J., Buongiorno, J. and Trancik, J. (2020) “Sources of Cost Overrun in Nuclear Power Plant Construction Call for a New Approach to Engineering Design”, Joule, 4(11).

Get It Right Initiative (2015) Strategy for Change; and (2016) Call to Action.

Goolsbee, A. and Syverson, C. (2023) The Strange and Awful Path of Productivity in the U.S. Construction Sector, NBER Working Paper 30845.

McKinsey Global Institute (2017) Reinventing Construction: A Route to Higher Productivity, February.

Behavioural science: modelling, demonstration and transfer

Ashford, D., Bennett, S. J. and Davids, K. (2006) “Observational modeling effects for movement dynamics and movement outcome measures across differing task constraints: a meta-analysis”, Journal of Motor Behavior, 38(3), pp. 185–205.

Bandura, A. (1986) Social Foundations of Thought and Action: A Social Cognitive Theory. Englewood Cliffs: Prentice-Hall.

Burke, M. J., Sarpy, S. A., Smith-Crowe, K., Chan-Serafin, S., Salvador, R. O. and Islam, G. (2006) “Relative effectiveness of worker safety and health training methods”, American Journal of Public Health, 96(2), pp. 315–324.

Christian, M. S., Bradley, J. C., Wallace, J. C. and Burke, M. J. (2009) “Workplace safety: a meta-analysis of the roles of person and situation factors”, Journal of Applied Psychology, 94(5), pp. 1103–1127.

Cook, D. A., Hatala, R., Brydges, R., et al. (2011) “Technology-enhanced simulation for health professions education: a systematic review and meta-analysis”, JAMA, 306(9), pp. 978–988.

Cook, D. A., Brydges, R., Hamstra, S. J., et al. (2012) “Comparative effectiveness of technology-enhanced simulation versus other instructional methods”, Simulation in Healthcare, 7(5), pp. 308–320.

Heyes, C. and Catmur, C. (2022) “What happened to mirror neurons?”, Perspectives on Psychological Science, 17(1), pp. 153–168.

Höffler, T. N. and Leutner, D. (2007) “Instructional animation versus static pictures: a meta-analysis”, Learning and Instruction, 17(6), pp. 722–738.

Taylor, P. J., Russ-Eft, D. F. and Chan, D. W. L. (2005) “A meta-analytic review of behavior modeling training”, Journal of Applied Psychology, 90(4), pp. 692–709.

Pipeline, workforce and technology

Associated General Contractors of America (2026) 2026 Construction Hiring and Business Outlook.

Infrastructure Australia (2025) 2025 Infrastructure Market Capacity Report, 13 November.

McKinsey & Company (2025) The infrastructure moment, 9 September.

Ministry of Housing, Communities and Local Government (2025) Housing supply: net additional dwellings, England, 2024–25, 20 November.

Zacua Ventures (2026) Construction Robotics Report 2026, March.

Recommended Posts

Interactive industry training solutions: eLearning using CGI

Technical Empathy in Operational Readiness: Train Driver and Train Controllers

Collaborative Development for 3D Real-Time Planning Models