← All Articles

The Hidden Risk of Incomplete AI Answers: Why AEC Cannot Afford Half-Truths

11 min readGovernance & RiskSharePDF

Listen to this article

The Hidden Risk of Incomplete AI Answers: Why AEC Cannot Afford Half-Truths

0:00
Jump to a section

Executive Summary

Generative AI’s most dangerous failure mode is a confident, incomplete answer, not a crash. In deterministic safety domains like AEC, that makes probabilistic AI unsuitable for direct use in safety-critical decisions, and liability for its outputs falls entirely on the deploying firm, not the system.

33-79%

hallucination rate OpenAI itself measured on its own reasoning models, depending on the benchmark

1 in 5

AI coding suggestions contain factual errors, even at 50-65% headline accuracy

18.9% → 91.5%

error-detection rate improvement from adding structured human-in-the-loop oversight

Core conclusions

  • AI systems bear no legal responsibility for their outputs: negligence claims attach entirely to the firm that deployed them, creating an accountability asymmetry vendor contracts rarely resolve.
  • The EU AI Act classifies AI used as a safety component in construction as “high-risk,” mandating lifecycle risk management, data governance, explainability, and continuous monitoring.
  • The right architecture uses generative AI only for exploratory, non-binding analysis, and requires deterministic, auditable, human-reviewed systems for every safety-critical decision.

The most dangerous moment in the deployment of artificial intelligence is not when systems fail catastrophically. It is when they deliver partial truths with complete confidence. In the architecture, engineering, and construction (AEC) sectors, this distinction matters profoundly. Incomplete information doesn’t just underperform. It kills.

Incomplete information doesn’t just underperform. It kills.

Why AI Produces Dangerous Incomplete Answers

Current large language models and generative AI systems operate through statistical pattern matching rather than true comprehension or reasoning. When confronted with an incomplete dataset, a novel scenario, or a knowledge gap, these systems don’t acknowledge absence. They fabricate. Model scaling doesn’t fix this. It’s inherent to how probabilistic AI functions.

Research demonstrates this consistently: the hallucination and coding-accuracy figures above hold across both general-purpose and task-specific models, and the two problems compound each other. How confidently wrong the answer sounds is the critical failure, more than the error rate itself. These systems don’t signal uncertainty. They present conjectures as certainties because the model optimises for fluency and coherence. The fix is architectural, not aspirational: building systems that can abstain when uncertain.

These systems don’t signal uncertainty. They present conjectures as certainties.

Construction and engineering workflows are not information-seeking conversations. They are deterministic domains where half-answers are half-failures. Incomplete structural load calculations mean the structure fails. Omitted soil conditions in geotechnical assessments mean foundations fail. Overlooked edge cases (seismic activity, thermal expansion, material degradation) produce irreversible operational and safety consequences. Generative AI systems are architected for probabilistic inference, generating varied outputs for identical inputs. This adaptive variability is celebrated in creative applications. In deterministic domains, it is a liability. The construction industry has recognised this risk and adopted a cautious “wait and see” approach precisely because AI lacks the reasoning capacity required for safety-critical applications.

Mindmap titled Consequences of Incomplete AEC Calculations, showing five branches from a central Half-Answers in AEC Domains node: Incomplete Structural Load Calculations leading to Structure Fails; Omitted Soil Conditions leading to Foundations Fail; and Overlooked Seismic Activity, Overlooked Thermal Expansion, and Overlooked Material Degradation each leading to Irreversible Safety Consequences
Three of the five branches land in the same place: “irreversible.” That’s the tell that these are the baseline a deterministic domain has to design for, not edge cases to patch later.

The Liability and Regulatory Reality

AI assumes no responsibility for its outputs. Responsibility remains with the operator, the deployer, and the engineering firm. Under existing law in most jurisdictions, construction companies have a legal duty of care. If an AI system is deployed without rigorous testing, validation, and human oversight, negligence claims attach to the firm deploying the system, not to the system itself. This creates an accountability asymmetry: the firm bears unlimited liability whilst the AI system bears none. The system cannot be held negligent. It cannot defend its reasoning. It cannot provide testimony in court about what information it considered or rejected. Contracts with AI vendors rarely resolve this issue adequately: most vendor agreements include clauses that disclaim liability, require indemnification from the deployer, and restrict the types of claims construction firms can bring.

Emerging regulatory frameworks are beginning to acknowledge this danger. The European Union AI Act imposes specific obligations on providers and deployers of “high-risk” AI systems, a classification that includes AI systems used as safety components in construction and built-environment systems. Requirements include an iterative risk management system throughout the entire AI system lifecycle, data governance with representative and systematically assessed training datasets, accuracy and robustness requirements with declared metrics, explainability enabling humans to understand and validate decision-making processes, and continuous monitoring of systems that continue learning after deployment. These requirements exist because regulators understand the fundamental mismatch between probabilistic AI and deterministic requirements.

Requirements for Safe Deployment

The alternative to probabilistic AI is a deterministic architecture: systems that produce identical outputs for identical inputs, operate according to explicitly defined rules, can be audited and traced, and explicitly acknowledge when they lack sufficient information to produce an answer. For AEC applications, structural design optimisation, load-bearing verification, material selection validation, and safety assessment must operate in this deterministic mode. They can be augmented by generative AI for exploratory phases (suggesting design variants, identifying novel material combinations, generating planning scenarios), but the moment the output enters the domain of safety-critical decision-making, it must transition to a deterministic, auditable, and human-reviewable space.

Before deploying any AI system in AEC contexts, organisations must establish explicit, verified, and enforceable safeguards. Structured human-in-the-loop oversight is non-negotiable. AI outputs do not deploy directly into construction or engineering decisions; qualified domain experts must review, validate, and either approve or reject each material output. Research shows structured human oversight increases error detection from 18.9% to 91.5%. Designing that oversight layer deliberately is the subject of Redesigning Oversight Architectures. Explicit accuracy and completeness criteria must be defined before deployment. Complete traceability and auditability throughout the lifecycle must be maintained, from data ingestion through model training, validation, deployment, and ongoing operation. Explainability must be tied to every material decision, accompanied by a technically defensible explanation understandable by domain engineers without AI expertise. Continuous monitoring and post-deployment assurance must be implemented, with deviations triggering investigation and corrective action. Finally, contracts with AI vendors must explicitly define who bears liability if the system fails, what warranties are provided, and what audit rights the deployer maintains over model training and data.

Mindmap of six safeguards required for AI deployment in AEC: Human-in-the-Loop Oversight with qualified domain experts and material output review; Accuracy and Completeness Criteria defined pre-deployment; Explainability with defensible explanations for domain engineers; Traceability and Auditability across the full lifecycle; Continuous Monitoring with deviation detection and corrective action; Vendor Liability Contracts defining liability and audit rights
None of these six are optional extras. Skip one and the accountability gap between what the AI produced and what the firm can defend in court stays open.

The Path to Accountability

The reason these requirements exist is not to impose a regulatory burden for its own sake. It is because the stakes are unforgiving. When AI systems fail in low-consequence environments, the cost is inconvenience and rework. When AI systems fail in AEC environments, the consequences can include injury, death, or catastrophic infrastructure failure. AI will not be held responsible for these outcomes. The operator will. The firm will. The engineers who deployed and supervised the system will face scrutiny, liability, and possibly criminal negligence charges if the failure was foreseeable and inadequately mitigated.

The path is neither “embrace AI without reservation” nor “abandon AI entirely.” It is: deploy AI in AEC only within a governance framework that acknowledges the domain’s deterministic safety requirements and establishes verification practices that exceed regulatory minimums. This means using probabilistic generative AI exclusively for exploratory, non-binding analysis; transitioning to deterministic systems for all safety-critical decisions; implementing structured human-in-the-loop validation for every material output; maintaining complete audit trails and explainability mechanisms; and allocating clear contractual and operational responsibility for failures.

The most dangerous phrase in AI-enabled engineering is “the system said so.”

That is not professional engineering judgment. That is an abdication of responsibility. The required phrase is: “We have reviewed the system’s analysis, considered the underlying data and logic, conducted independent verification, and determined this conclusion is sound.”

The guard cannot be lowered. It must be installed, tested, documented, and continuously monitored. AI is a powerful tool. But in AEC, it is a tool that requires more, not less, expert human oversight.

AI is a powerful tool. But in AEC, it is a tool that requires more, not less, expert human oversight.


Evidence & Methodology

The oversight number in this piece comes from clinical research, not construction. I am borrowing it because it is the closest real measurement I have found, and you should know that before you cite it in a safety case.

ClaimSourceGrade
OpenAI’s own reasoning models hallucinate 33-79% of the time depending on the benchmarkOpenAI’s own o3 and o4-mini system cardMeasured
Structured human oversight raised error detection from 18.9% to 91.5%Umakanth, 2025, a study of human-in-the-loop validation in clinical GenAI settings, not constructionAdjacent study
The EU AI Act classifies AI safety components in construction as “high-risk,” with lifecycle risk management and data-governance dutiesThe EU AI Act itself, Articles 6, 10, and 13Legal text
1 in 5 AI coding suggestions contain factual errorsAugment Code, a vendor whose product competes in this spaceVendor source

Sources

  1. OpenAI. (2025). “OpenAI o3 and o4-mini System Card,” documenting hallucination rates of 33-79% across benchmarks.
  2. Augment Code. (2025). “Deterministic AI for Predictable Coding,“
  3. AI21 Labs. (2025). “What are AI Hallucinations? Signs, Risks, & Prevention,“
  4. Kubiya AI. (2025). “What Is Deterministic AI?”
  5. Graitec. (2024). “AI Meets Sustainability: Innovations in AEC and MFG,“
  6. European Union. (2024). EU AI Act, Article 6 & Annex III, classification rules for high-risk AI systems.
  7. European Union. (2024). EU AI Act, Article 10, data governance requirements.
  8. European Union. (2024). EU AI Act, Article 13, transparency and post-market monitoring requirements.
  9. Umakanth, A. A. M. (2025). “Human-in-the-Loop Architectures for Validating GenAI Outputs in Clinical Settings,” European Journal of Computer Science and Information Technology.

Free tool

ISO 42001 AIMS Readiness Checklist

Check your AEC deployment against the 36 clauses the EU AI Act effectively requires for any high-risk safety-component AI system.

Was this useful?

Terence Kok
Before You Go

I spent years around engineers who'd never sign a drawing they hadn't checked twice, so watching teams treat an AI's confident answer like a stamped calculation worried me enough to write this down. The 33 to 79 percent hallucination range sticks with me because that's the distance between a model sounding right and a beam holding weight. This one is for the engineer who's been told to move faster and isn't sure where the line is anymore. Draw it clearly, and you'll sleep better.

Terence Kok