PofoliaShared via Pofolia

ISPRS International Journal of Geo-Information· 2026Q1

Empirical Post-Generation Correctness Assessment for Natural-Language Access to Urban PostGIS Databases

Jiqiu Deng, C L Zhang, Hui Zhang, Longbo Li et al.

Short summary

Spatial-safe improves the correctness of natural-language queries to urban databases by 0.062 AUROC on average across nine transfer directions in a controlled benchmark, raising coverage from 0.457 to 0.669.

AI-generated from the title and abstract; the full text is not read.

Key points

  • Spatial-safe improved executed non-empty AUROC by 0.062 on average across nine transfer directions in a controlled benchmark.
  • Mean conditional target coverage within the executed non-empty operating subset rose from 0.457 to 0.669.
  • The current benchmark, with only 77 independent SQL-skeleton groups, is insufficient for adequately powered structure-disjoint performance comparisons.
  • The findings support correctness-ranking improvement in controlled settings, not universal structure-disjoint generalization or formal verification.

AI-generated from the title and abstract; the full text is not read.

Abstract

Natural-language interfaces make urban spatial databases accessible, but generated structured query language (SQL) can execute successfully while returning a semantically incorrect answer. We evaluate a post-generation correctness-ranking layer on 18,900 candidates from Nanjing, Wuhan, and Shenzhen. The original city split is a controlled, highly template-aligned benchmark rather than a structure-disjoint deployment test. Under strict source-side separation of fitting, isotonic calibration, threshold selection, and target evaluation, Spatial-safe improves executed non-empty area under the receiver operating characteristic curve (AUROC) in all nine transfer directions by 0.062 on average. Grouped-threshold area under the risk–coverage curve (AURC) is favorable in 8/9 directions, and calibration is mixed. Source-selected thresholds raise mean conditional target coverage within the executed non-empty operating subset from 0.457 to 0.669, while mean empirical risk rises from 0.110 to 0.136. Unseen-template and SQL-skeleton-disjoint effects are smaller and mixed. To determine whether the existing benchmark could support the adequately powered structure-disjoint comparison, we applied a predeclared power/data-sufficiency feasibility gate. With only 77 independent SQL-skeleton groups, 42/60 required cells fail the gate; therefore, the present benchmark cannot support an adequately powered structure-disjoint performance comparison without additional skeleton-diverse questions. We do not substitute an underpowered estimate for that missing evidence. The evidence therefore supports empirical correctness-ranking improvement in this controlled benchmark, not universal structure-disjoint generalization, formal SQL verification, benchmark-wide human validation of all correctness labels, or target-domain risk guarantees.

The authors' abstract, as published at the source. ISPRS International Journal of Geo-Information, 2026 · DOI ↗

TakeawaysPremium
Ask the paperFree account

Continue with a free account

Ask the paper: 3 free questions a day about this paper; save it, get its citation, new summaries every day for your field. Takeaways are Premium.

Continue free on the web

Sign in with Google or Apple; no card needed. You come back to this paper.

On your phone:

Field: Geography, Planning and Development

Geography, Planning and DevelopmentSocial Sciences