PofoliaShared via Pofolia

Proceedings of the ACM on Interactive Mobile Wearable and Ubiquitous Technologies· 2026Q1

Feasibility of Using a Multi-Agent LLM System to Correct Annotations and Support Low-Effort Activity Labeling

Ha Le, Akshat Choube, Varun Mishra, STEPHEN S. INTILLE

Short summary

A multi-agent LLM system, GLOSS4HAR, improves activity annotation quality by up to 9.9% F1 score and reconstructs activity timelines with 75-92% F1 accuracy, by reconciling self-reports and passive sensing data.

AI-generated from the title and abstract; the full text is not read.

Key points

  • Introduced GLOSS4HAR, a multi-agent LLM system for correcting activity annotations and supporting low-effort labeling.
  • GLOSS4HAR improves annotation quality by up to 9.9% in F1 score.
  • Reconstructs activity timelines with 75-92% F1 score by combining passive sensing and self-reports.
  • Mimics human sensemaking to reconcile inconsistencies between data sources.

AI-generated from the title and abstract; the full text is not read.

Abstract

Accurate measurement of behaviors is critical for research in human-computer interaction, ubiquitous computing, and personal health informatics, because it underpins many health tracking and intervention systems. Most data collection studies, however, still rely on participants' self-reports and manual annotations, which often under- or over-estimate activity duration and type, and require substantial effort from researchers to clean and validate. An automated system that can combine passive sensing data with participants' self-reports might detect inconsistencies and suggest corrections. We introduce GLOSS4HAR, a multi-agent LLM-based system designed to mimic human sensemaking and assist researchers in cleaning and refining activity annotations. We demonstrate the potential of GLOSS4HAR in two key tasks: (1) correcting and reconciling participant self-annotations, and (2) triangulating passive sensing data with different forms of lightweight self-reports to generate accurate activity timelines. Our evaluation shows that GLOSS4HAR improves annotation quality by up to 9.9% in F1 score and can reconstruct activity timelines that align with human annotations at 75-92% F1. Based on our findings, we discuss the implications of our work for the next generation of activity annotation systems that might use human-AI collaboration.

The authors' abstract, as published at the source. Proceedings of the ACM on Interactive Mobile Wearable and Ubiquitous Technologies, 2026 · DOI ↗

TakeawaysPremium
Ask the paperFree account

Continue with a free account

Ask the paper: 3 free questions a day about this paper; save it, get its citation, new summaries every day for your field. Takeaways are Premium.

Continue free on the web

Sign in with Google or Apple; no card needed. You come back to this paper.

On your phone:

Field: Information Systems and Management

Information Systems and ManagementDecision Sciences