Listen to this article
Executive Summary
Generative AI’s most dangerous failure mode isn’t a crash. It’s a confident, incomplete answer. In deterministic safety domains like AEC, that makes probabilistic AI unsuitable for direct use in safety-critical decisions, and liability for its outputs falls entirely on the deploying firm, not the system.
33-79%
hallucination rate for ChatGPT-class models, depending on task complexity
1 in 5
AI coding suggestions contain factual errors, even at 50-65% headline accuracy
18.9% → 91.5%
error-detection rate improvement from adding structured human-in-the-loop oversight
Core conclusions
- AI systems bear no legal responsibility for their outputs: negligence claims attach entirely to the firm that deployed them, creating an accountability asymmetry vendor contracts rarely resolve.
- The EU AI Act classifies AI used as a safety component in construction as “high-risk,” mandating lifecycle risk management, data governance, explainability, and continuous monitoring.
- The right architecture uses generative AI only for exploratory, non-binding analysis, and requires deterministic, auditable, human-reviewed systems for every safety-critical decision.
The risk, in ten slides
Save it, share it, or send it to whoever is signing off your AI deployment.










The most dangerous moment in the deployment of artificial intelligence is not when systems fail catastrophically. It is when they deliver partial truths with complete confidence. In the architecture, engineering, and construction (AEC) sectors, this distinction matters profoundly. Incomplete information doesn’t just underperform. It kills.
Incomplete information doesn’t just underperform. It kills.
Why AI Produces Dangerous Incomplete Answers
Current large language models and generative AI systems operate through statistical pattern matching rather than true comprehension or reasoning. When confronted with an incomplete dataset, a novel scenario, or a knowledge gap, these systems don’t acknowledge absence. They fabricate. This isn’t a bug that can be fixed by model scaling. It’s inherent to how probabilistic AI functions.
Research demonstrates this consistently: ChatGPT-class models hallucinate between 33% and 79% of the time depending on task complexity, and AI coding assistants maintain accuracy rates between 50–65%, with approximately 1 in 5 suggestions containing factual errors. The critical failure isn’t the error rate itself. It’s the confidence interval for incorrect answers. These systems don’t signal uncertainty. They present conjectures as certainties because the model optimises for fluency and coherence rather than factual verification. The fix is architectural, not aspirational: building systems that can actually abstain when uncertain, rather than guessing with false confidence.
These systems don’t signal uncertainty. They present conjectures as certainties.
Construction and engineering workflows are not information-seeking conversations. They are deterministic domains where half-answers are half-failures. Incomplete structural load calculations mean the structure fails. Omitted soil conditions in geotechnical assessments mean foundations fail. Overlooked edge cases (seismic activity, thermal expansion, material degradation) produce irreversible operational and safety consequences. Generative AI systems are architected for probabilistic inference, generating varied outputs for identical inputs. This adaptive variability is celebrated in creative applications. In deterministic domains, it is a liability. The construction industry has recognised this risk and adopted a cautious “wait and see” approach precisely because AI lacks the reasoning capacity required for safety-critical applications.
The Liability and Regulatory Reality
AI assumes no responsibility for its outputs. Responsibility remains with the operator, the deployer, and the engineering firm. Under existing law in most jurisdictions, construction companies have a legal duty of care. If an AI system is deployed without rigorous testing, validation, and human oversight, negligence claims attach to the firm deploying the system, not to the system itself. This creates an accountability asymmetry: the firm bears unlimited liability whilst the AI system bears none. The system cannot be held negligent. It cannot defend its reasoning. It cannot provide testimony in court about what information it considered or rejected. Contracts with AI vendors rarely resolve this issue adequately: most vendor agreements include clauses that disclaim liability, require indemnification from the deployer, and restrict the types of claims construction firms can bring.
Emerging regulatory frameworks are beginning to acknowledge this danger. The European Union AI Act imposes specific obligations on providers and deployers of “high-risk” AI systems, a classification that includes AI systems used as safety components in construction and built-environment systems. Requirements include an iterative risk management system throughout the entire AI system lifecycle, data governance with representative and systematically assessed training datasets, accuracy and robustness requirements with declared metrics, explainability enabling humans to understand and validate decision-making processes, and continuous monitoring of systems that continue learning after deployment. These requirements exist because regulators understand the fundamental mismatch between probabilistic AI and deterministic requirements.
Requirements for Safe Deployment
The alternative to probabilistic AI is a deterministic architecture: systems that produce identical outputs for identical inputs, operate according to explicitly defined rules, can be audited and traced, and explicitly acknowledge when they lack sufficient information to produce an answer. For AEC applications, structural design optimisation, load-bearing verification, material selection validation, and safety assessment must operate in this deterministic mode. They can be augmented by generative AI for exploratory phases (suggesting design variants, identifying novel material combinations, generating planning scenarios), but the moment the output enters the domain of safety-critical decision-making, it must transition to a deterministic, auditable, and human-reviewable space.
Before deploying any AI system in AEC contexts, organisations must establish explicit, verified, and enforceable safeguards. Structured human-in-the-loop oversight is non-negotiable. AI outputs do not deploy directly into construction or engineering decisions; qualified domain experts must review, validate, and either approve or reject each material output. Research shows structured human oversight increases error detection from 18.9% to 91.5%. Designing that oversight layer deliberately, rather than bolting it on, is the subject of Redesigning Oversight Architectures. Explicit accuracy and completeness criteria must be defined before deployment. Complete traceability and auditability throughout the lifecycle must be maintained, from data ingestion through model training, validation, deployment, and ongoing operation. Explainability must be tied to every material decision, accompanied by a technically defensible explanation understandable by domain engineers without AI expertise. Continuous monitoring and post-deployment assurance must be implemented, with deviations triggering investigation and corrective action. Finally, contracts with AI vendors must explicitly define who bears liability if the system fails, what warranties are provided, and what audit rights the deployer maintains over model training and data.
The Path to Accountability
The reason these requirements exist is not to impose a regulatory burden for its own sake. It is because the stakes are genuinely unforgiving. When AI systems fail in low-consequence environments, the cost is inconvenience and rework. When AI systems fail in AEC environments, the consequences can include injury, death, or catastrophic infrastructure failure. AI will not be held responsible for these outcomes. The operator will. The firm will. The engineers who deployed and supervised the system will face scrutiny, liability, and possibly criminal negligence charges if the failure was foreseeable and inadequately mitigated.
The path is neither “embrace AI without reservation” nor “abandon AI entirely.” It is: deploy AI in AEC only within a governance framework that acknowledges the domain’s deterministic safety requirements and establishes verification practices that exceed regulatory minimums. This means using probabilistic generative AI exclusively for exploratory, non-binding analysis; transitioning to deterministic systems for all safety-critical decisions; implementing structured human-in-the-loop validation for every material output; maintaining complete audit trails and explainability mechanisms; and allocating clear contractual and operational responsibility for failures.
The most dangerous phrase in AI-enabled engineering is “the system said so.”
That is not professional engineering judgment. That is an abdication of responsibility. The required phrase is: “We have reviewed the system’s analysis, considered the underlying data and logic, conducted independent verification, and determined this conclusion is sound.”
The guard cannot be lowered. It must be installed, tested, documented, and continuously monitored. AI is a powerful tool. But in AEC, it is a tool that requires more, not less, expert human oversight.
AI is a powerful tool. But in AEC, it is a tool that requires more, not less, expert human oversight.
References:
Talkspace (2025). The Dangers of ChatGPT Hallucinations. OpenAI testing documented hallucination rates of 33-79%.
Augmentcode (2025). Deterministic AI for Predictable Coding. AI coding assistants achieve 50-65% accuracy with 1 in 5 suggestions containing factual errors.
AI21 Labs (2025). Specific risks of AI hallucinations. LLMs optimised for fluent language rather than factual verification.
Kubiya AI (2025). Deterministic AI Architecture: Why It Matters. Deterministic systems produce identical outputs for identical inputs; probabilistic systems generate varied responses.
Graitec (2024). AI Meets Sustainability: Innovations in AEC and MFG. Construction firms adopting cautious “wait and see” approach due to unreliability risks.
Construction Legal Services (2025). AI and the Construction Industry. Liability question complex; firms need contractual protections and clear responsibility definitions.
Insulation.org (2025). Legal Risks of AI in Construction. Negligence claims possible if insufficient testing or blind reliance; strict liability applies if deployed without adequate validation.
Keysight Technologies (2026). Building Trust Into AI for Safety Critical Systems. EU AI Act and ISO PAS 8800 mandate transparency, traceability, and risk-based validation.
European Commission (2024). Classification Rules for High-Risk AI Systems. Risk management system mandatory throughout entire lifecycle.
EU AI Act (2024). Article 5-7. Data governance requirements: representative, error-free datasets with bias assessment.
EU AI Act (2024). Article 10-15. Accuracy, robustness, cybersecurity requirements with declared metrics and consistent performance throughout lifecycle.
Port.io (2024). What is Deterministic AI? Deterministic systems operate on predefined rules; identical inputs produce identical outputs.
eajournals (2025). Human-in-the-Loop Architectures for Validating GenAI. Structured HITL validation increased error detection from 18.9% to 91.5% in medical AI.
Journals MRI India (2025). Explainable AI for Critical Infrastructure Monitoring and Control. Feature attribution methods (SHAP, LIME, counterfactual explanations) essential for transparency.
EU AI Act (2024). Article 13(e). Post-market monitoring and continuous feedback loop management mandatory for systems that learn after deployment.
Free tool
ISO 42001 AIMS Readiness Checklist
Check your AEC deployment against the 36 clauses the EU AI Act effectively requires for any high-risk safety-component AI system.
Was this useful?
