ISPRS International Journal of Geo-Information· 2026Q1
Empirical Post-Generation Correctness Assessment for Natural-Language Access to Urban PostGIS Databases
- 0citations
- Q1SCImago
- 2026year
Short summary
Spatial-safe improves the correctness of natural-language queries to urban databases by 0.062 AUROC on average across nine transfer directions in a controlled benchmark, raising coverage from 0.457 to 0.669.
AI-generated from the title and abstract; the full text is not read.
Key points
- Spatial-safe improved executed non-empty AUROC by 0.062 on average across nine transfer directions in a controlled benchmark.
- Mean conditional target coverage within the executed non-empty operating subset rose from 0.457 to 0.669.
- The current benchmark, with only 77 independent SQL-skeleton groups, is insufficient for adequately powered structure-disjoint performance comparisons.
- The findings support correctness-ranking improvement in controlled settings, not universal structure-disjoint generalization or formal verification.
AI-generated from the title and abstract; the full text is not read.
Abstract
Natural-language interfaces make urban spatial databases accessible, but generated structured query language (SQL) can execute successfully while returning a semantically incorrect answer. We evaluate a post-generation correctness-ranking layer on 18,900 candidates from Nanjing, Wuhan, and Shenzhen. The original city split is a controlled, highly template-aligned benchmark rather than a structure-disjoint deployment test. Under strict source-side separation of fitting, isotonic calibration, threshold selection, and target evaluation, Spatial-safe improves executed non-empty area under the receiver operating characteristic curve (AUROC) in all nine transfer directions by 0.062 on average. Grouped-threshold area under the risk–coverage curve (AURC) is favorable in 8/9 directions, and calibration is mixed. Source-selected thresholds raise mean conditional target coverage within the executed non-empty operating subset from 0.457 to 0.669, while mean empirical risk rises from 0.110 to 0.136. Unseen-template and SQL-skeleton-disjoint effects are smaller and mixed. To determine whether the existing benchmark could support the adequately powered structure-disjoint comparison, we applied a predeclared power/data-sufficiency feasibility gate. With only 77 independent SQL-skeleton groups, 42/60 required cells fail the gate; therefore, the present benchmark cannot support an adequately powered structure-disjoint performance comparison without additional skeleton-diverse questions. We do not substitute an underpowered estimate for that missing evidence. The evidence therefore supports empirical correctness-ranking improvement in this controlled benchmark, not universal structure-disjoint generalization, formal SQL verification, benchmark-wide human validation of all correctness labels, or target-domain risk guarantees.
The authors' abstract, as published at the source. ISPRS International Journal of Geo-Information, 2026 · DOI ↗
Continue with a free account
Ask the paper: 3 free questions a day about this paper; save it, get its citation, new summaries every day for your field. Takeaways are Premium.
Continue free on the webSign in with Google or Apple; no card needed. You come back to this paper.
On your phone:
Field: Geography, Planning and Development
Geography, Planning and DevelopmentSocial Sciences