NTCIR-19

NTCIR19-Lifelog Task

Advancing Lifelog Analytics and Retrieval

Evaluation

The NTCIR19-Lifelog task employs a comprehensive evaluation methodology to assess the performance of participating systems across all sub-tasks. We will use post-evaluation as the primary evaluation approach, which involves assessing system performance after all submissions have been received, allowing for thorough analysis and comparison of results.

General Evaluation Metrics

The evaluation will utilize multiple metrics to comprehensively assess system performance:

  • Precision and Recall: Standard information retrieval metrics to measure the accuracy and completeness of retrieved results.
  • Mean Average Precision (MAP): A single-value metric that summarizes the precision-recall curve, providing an overall measure of retrieval quality.
  • Normalized Discounted Cumulative Gain (NDCG): A ranking-based metric that considers the position of relevant items in the result list.
  • Success at K (S@K): Measures whether at least one relevant item appears in the top K results.
  • Semantic Similarity Scores: For tasks involving semantic understanding, we evaluate the semantic relevance of retrieved items to the query.

Lifelog Task Evaluation

LSAT Evaluation

For the Lifelog Semantic Access Task (LSAT), the evaluation follows the same methodology as NTCIR18-Lifelog6:

Evaluation Tool: The trec-eval programme is employed to generate result scores for each run.

Relevance Judgments: Binary relevance judgments are created through a pooled approach, where human assessors manually evaluate each submitted image for each topic, up to a maximum of 100 images per topic, per run, per participant. The relevance judgments are binary and are informed by the immediate context of each image when important for understanding the semantic relevance.

Evaluation Process:

  1. All participant runs are collected by the submission deadline.
  2. Expert assessors prepare relevance judgments using the pooled approach.
  3. The trec-eval programme computes evaluation metrics automatically.
  4. Selected results undergo manual review to ensure quality.
  5. Statistical analysis is performed to compare systems.
  6. Results are distributed to participants for analysis.

Topic Types: There are two types of topics:

  • ADHOC: Topics that may have many moments in the collection that are relevant.
  • KNOWN-ITEM: Topics with one (or few) relevant moments in the collection.

CASTLE Task Evaluation

CSAT Evaluation

The CASTLE Semantic Access Task (CSAT) evaluation focuses on assessing the retrieval of key interactions or events from multimodal collaborative session data.

Evaluation Approach: Each submitted result is a single occurrence of the queried event, reported as a (day, time range, source) triple. The relevant set for each topic is an open set, and there is no predefined shot segmentation: assessors judge each submitted time range directly.

Evaluation Metrics: Scoring is based on the fraction of correct entries (precision) together with standard information-retrieval metrics computed over the score-ranked results, primarily mean Average Precision (mAP) and recall@k. nDCG and Success@K may also be reported for completeness. Because the relevant set is open and manually judged, recall is measured relative to the pooled set of correct instances found across all runs.

CAST-Seg Evaluation

CAST-Seg has no pre-defined ground truth, so we do not use Pk, WindowDiff, boundary F1, or a fixed temporal tolerance. Submitted segmentations are reviewed by expert assessors against the CASTLE 2024 recordings and transcripts for boundary plausibility, semantic coherence, label appropriateness, and coverage across each session.

Runs that pass review will be pooled and released as a community benchmark, with attribution to contributing teams. See the Formal Run and Submission pages for task and file-format details.

Recipe Generation Evaluation

Recipe Generation evaluates whether systems can reconstruct plausible, session-grounded recipes for dishes prepared during the CASTLE collaborative cooking sessions. Like CAST-Seg, this task has no single canonical ground-truth recipe file for scoring: participants cooked from personal notes, group-scale ingredient lists, and on-the-day adaptations, so the correct output is best judged by whether a generated recipe is faithful to what actually happened in the recordings rather than by exact string match to a reference document.

Assessment criteria: For each target dish, assessors review the submitted recipe against the CASTLE multimodal recordings (video, audio, transcripts), the menu schedule, and the Castle Cookbook reference materials. Submissions are judged on:

  • Relevance and fidelity: Do the listed ingredients and steps correspond to the dish prepared on the scheduled day, as observed in the session data?
  • Coherence and completeness: Is the procedure logically ordered and sufficiently detailed to reproduce the dish, without major missing stages?
  • Multimodal grounding: Does the recipe reflect information that could reasonably be inferred from the collaborative session (preparation actions, timing, group cooking context), not only generic cookbook text?
  • Plausibility: Are quantities, techniques, and ingredient choices realistic for the CASTLE kitchen setting?

Detailed dish targets and supporting resources (menu schedule, Castle Cookbook) are described on the Formal Run page; there is no separate topic-release package beyond those materials and the CASTLE dataset itself.

Evaluation Criteria

Systems are evaluated based on several criteria:

  • Accuracy: The correctness of retrieved items in relation to the query.
  • Completeness: The ability to retrieve all relevant items for a given query.
  • Efficiency: The computational efficiency and response time of the system.
  • Robustness: The system's performance across different types of queries and data conditions.
  • Novelty: The innovation and contribution of the approach to the field.

Evaluation Timeline

The evaluation timeline aligns with the task schedule:

  • Formal Run Phase: April–August 2026 - Participants submit their formal runs during this period.
  • Evaluation Period: September 2026 - Post-evaluation is conducted and draft overview paper is prepared.
  • Results Release: October 2026 - Evaluation results are returned to participants.
  • Camera-Ready Papers: November 2026 - Participants submit final camera-ready papers.
  • Conference: December 2026 - NTCIR-19 Conference in Tokyo, Japan.

Fairness and Reproducibility

To ensure fairness and reproducibility:

  • LSAT and CSAT runs are scored with documented evaluation scripts once relevance judgments are complete.
  • CAST-Seg and Recipe Generation are assessed with published criteria applied consistently by expert reviewers.
  • Evaluation procedures and, where applicable, released benchmark materials are documented and made available to participants.
  • Participants are encouraged to provide detailed descriptions of their methods for reproducibility.
  • Any issues or discrepancies are addressed through a transparent review process.