Journal of Legal Affairs and Dispute Resolution in Engineering and Construction· 2026Q2
Evaluation of Prompt Engineering for LLM-Based Construction Contract Risk Analysis
- 0citations
- Q2SCImago
- 2026year
Short summary
Guided, clause-anchored prompting of GPT-4.1 significantly improved the accuracy, completeness, and legal grounding of LLM-generated risk analyses for construction contracts compared to less structured methods.
AI-generated from the title and abstract; the full text is not read.
Key points
- Guided, clause-anchored prompting improved LLM risk analysis accuracy, completeness, clarity, and legal salience for NEC4-style clauses.
- Less structured prompting strategies led to narrative drift and weaker linkage to contract clauses.
- GPT-4.1 was used in a controlled, human-in-the-loop evaluation with expert-derived reference risk lists.
- The clause-guided human-in-the-loop (CG-HITL) framework is presented as a conceptual approach for structured LLM use in contract risk assessment.
AI-generated from the title and abstract; the full text is not read.
Abstract
Abstract Large language models (LLMs) show increasing potential for complex analytical tasks; however, their reliability in high-stakes domains such as construction contract interpretation remains insufficiently understood. This study presents a controlled, human-in-the-loop evaluation of how prompt-engineering strategies influence the quality and perceived reliability of LLM-generated risk analyses for New Engineering Contract (NEC4)–style clauses. Generative pretrained transformer GPT-4.1 was applied in a deterministic configuration to identify contractor risks across six synthetic NEC4-style Z-clauses using five prompt strategies: baseline, direct, role based, chain of thought, and guided prompting. Thirty outputs were evaluated under a blinded protocol against expert-derived reference risk lists by a mixed-expertise cohort, supporting exploratory comparison. Assessment combined descriptive quantitative analysis of Likert-scale ratings of accuracy, completeness, clarity, legal salience, and hallucination risk with qualitative thematic analysis. Results indicate that guided, clause-anchored prompting produced more structured, legally grounded, and traceable outputs, whereas less constrained strategies exhibited narrative drift and weaker clause linkage. These findings suggest that structured prompting may support more consistent clause-level interpretation under controlled conditions, although conclusions remain bounded by the exploratory design, limited sample, and synthetic dataset. The study contributes to understanding how prompt design shapes interpretability and verifiability in contract analysis. The clause-guided human-in-the-loop (CG-HITL) framework is presented as a conceptual approach to support structured use of LLMs in preaward contract risk assessment.
The authors' abstract, as published at the source. Journal of Legal Affairs and Dispute Resolution in Engineering and Construction, 2026 · DOI ↗
Continue with a free account
Ask the paper: 3 free questions a day about this paper; save it, get its citation, new summaries every day for your field. Takeaways are Premium.
Continue free on the webSign in with Google or Apple; no card needed. You come back to this paper.
On your phone:
Field: Management Science and Operations Research
Management Science and Operations ResearchDecision Sciences