Change search
Link to record
Permanent link

Direct link
Publications (4 of 4) Show all publications
Westin, F. & Eklund, J. (2026). Evaluating LLM-Based Temporal Subject Assignment in Fiction: The Impact of Evaluation Frameworks on Metadata Interpretation. Cataloging & Classification Quarterly
Open this publication in new window or tab >>Evaluating LLM-Based Temporal Subject Assignment in Fiction: The Impact of Evaluation Frameworks on Metadata Interpretation
2026 (English)In: Cataloging & Classification Quarterly, ISSN 0163-9374, E-ISSN 1544-4554Article in journal (Refereed) Epub ahead of print
Abstract [en]

This study examines how automatically generated temporal subject metadata for fiction can be evaluated when time is represented as year intervals. Using 45 Swedish novels with manually assigned chronological headings as the basis for comparison with model-generated intervals, the study contrasts two evaluation perspectives: overlap-based classification (precision, recall, F1) and boundary-based deviation (RMSE). The analysis shows how different evaluation frameworks highlight different aspects of temporal alignment between predicted and manually assigned intervals. By combining these perspectives, the study makes visible different types of temporal discrepancies and shows how evaluation design influences what is counted as alignment in interval-based subject metadata. 

Place, publisher, year, edition, pages
Taylor & Francis Group, 2026
Keywords
Temporal subject assignment, large language model, evaluation framework, subject analysis, temporal metadata
National Category
Information Studies Information Systems, Social aspects
Research subject
Library and Information Science
Identifiers
urn:nbn:se:hb:diva-36048 (URN)10.1080/01639374.2026.2711280 (DOI)001850353100001 ()2-s2.0-105047392563 (Scopus ID)
Funder
University of BoråsUniversity of Borås
Available from: 2026-08-31 Created: 2026-08-31 Last updated: 2026-09-01Bibliographically approved
Westin, F. (2025). Time, technique and text: scoping review of temporal information extraction and categorisation in documents. Journal of Documentation, 81(7), 135-156
Open this publication in new window or tab >>Time, technique and text: scoping review of temporal information extraction and categorisation in documents
2025 (English)In: Journal of Documentation, ISSN 0022-0418, E-ISSN 1758-7379, Vol. 81, no 7, p. 135-156Article in journal (Refereed) Published
Abstract [en]

Purpose:

This paper presents an investigation of the concept of “time as aboutness” in various texts, including news articles, social media posts and historical documents. The purpose of this paper is to analyse different forms of temporal information and map the techniques used to extract and categorise this information.

Design/methodology/approach:

A scoping review method was adopted to analyse the chosen literature set. This approach allowed for an overview of the different text document types, the techniques used and their temporal information.

Findings:

The findings reveal six temporal types of time-related data analysis: social events, socio-political events, news events, temporal expressions, historical events and time periods. Studies analysing social media, news articles, Wikipedia entries and historical documents provide insights into event detection and categorisation. In these documents, time appears as sequences of events, temporal expressions or distinct periods. In news articles, time appears as a series of occurrences, while temporal expressions reveal how time is linguistically articulated and perceived. The analysis also covers event categorisation methods, emphasising machine learning techniques, natural language processing, large language models and rule-based systems.

Originality/value:

The analysis of different types of time and methods of extracting temporal information from various texts contributes original insights to the understanding of temporal information. The findings reveal a need for expanding document variety, particularly to include fiction literature and point to the potential use of language models for future temporal information categorisation.

Place, publisher, year, edition, pages
Emerald Group Publishing Limited, 2025
Keywords
Automated methods, Classification, Ctegorisation, Event, Temporal information, Time as aboutness
National Category
Information Studies
Research subject
Library and Information Science
Identifiers
urn:nbn:se:hb:diva-33459 (URN)10.1108/jd-11-2024-0267 (DOI)001456767100001 ()2-s2.0-105001519323 (Scopus ID)
Available from: 2025-04-17 Created: 2025-04-17 Last updated: 2026-01-30Bibliographically approved
Westin, F. (2024). Comparing Feature Engineering Techniques for the Time Period Categorisation of Novels. KO KNOWLEDGE ORGANIZATION, 51(5), 330-339
Open this publication in new window or tab >>Comparing Feature Engineering Techniques for the Time Period Categorisation of Novels
2024 (English)In: KO KNOWLEDGE ORGANIZATION, ISSN 0943-7444, Vol. 51, no 5, p. 330-339Article in journal (Refereed) Published
Abstract [en]

The growing number of literary works being produced and published has emphasised the importance of better cataloguing methods to handle the increasing volume effectively. One specific issue is the lack of organising works by time periods, which is crucial for understanding and organising literature. In this study, "time" refers to when the story's events occur or the narrative's temporal setting, like specific historical periods or events, rather than the publication date. Categorising literary works based on their historical settings can significantly improve accessibility for library patrons navigating online catalogues. However, time period categorisation is uncommon, primarily due to the resource-intensive nature of the process, which necessitates extensive analysis by librarians and cataloguers. To address this issue, this paper proposes evaluating different machine learning workflows to predict time periods for novels. The workflow comprises preprocessing, feature engineering, classification, and evaluation. The feature engineering techniques used are Latent Dirichlet Allocation (LDA), Word Embedding with Sentence-BERT (WE SBERT), and Term Frequency-Inverse Document Frequency (TF-IDF), and the classification algorithm used is Logistic Regression. The models are assessed using the F1 score, precision, and recall metrics. The time period categories used are Medieval, Era of Great Power, Age of Liberty, and Gustavian periods. The objective is to determine how effectively each model categorises Swedish historical fiction novels into their appropriate time period categories. By leveraging machine learning techniques, the research seeks to supplement the time period categorisation process, aiding cataloguers and ultimately enhancing the accessibility and usability of library collections.

Keywords
literary categorization, machine learning techniques, time period categorization
National Category
Information Systems
Identifiers
urn:nbn:se:hb:diva-32621 (URN)10.5771/0943-7444-2024-5 (DOI)001309320900006 ()
Available from: 2024-09-25 Created: 2024-09-25 Last updated: 2025-09-24Bibliographically approved
Westin, F. (2024). Time Period Categorization in Fiction: A Comparative Analysis of Machine Learning Techniques. Cataloging & Classification Quarterly, 1-30
Open this publication in new window or tab >>Time Period Categorization in Fiction: A Comparative Analysis of Machine Learning Techniques
2024 (English)In: Cataloging & Classification Quarterly, ISSN 0163-9374, E-ISSN 1544-4554, p. 1-30Article in journal (Refereed) Published
Abstract [en]

This study investigates the automatic categorization of time period metadata in fiction, a critical but often overlooked aspect of cataloging. Using a comparative analysis approach, the performance of three machine learning techniques, namely Latent Dirichlet Allocation (LDA), Sentence-BERT (SBERT), and Term Frequency-Inverse Document Frequency (TF-IDF) were assessed, by examining their precision, recall, F1 scores, and confusion matrix results. LDA identifies underlying topics within the text, TF-IDF measures word importance, and SBERT measures sentence semantic similarity. Based on F1-score analysis and confusion matrix outcomes, TF-IDF and LDA effectively categorize text data by time period, while SBERT performed poorly across all time period categories.

Keywords
Cataloging for digital resources; fiction, LDA, machine learning, SBERT, text analysis, TF-IDF, time period categorization
National Category
Computer and Information Sciences Information Studies
Research subject
Library and Information Science; Library and Information Science
Identifiers
urn:nbn:se:hb:diva-31755 (URN)10.1080/01639374.2024.2315548 (DOI)001189786900001 ()2-s2.0-85189510207 (Scopus ID)
Available from: 2024-04-15 Created: 2024-04-15 Last updated: 2025-09-24Bibliographically approved
Organisations
Identifiers
ORCID iD: ORCID iD iconorcid.org/0009-0003-6165-2184

Search in DiVA

Show all publications