Research vision

Language as a foundation for trustworthy AI.

Language mediates how modern AI systems represent and retrieve knowledge, use tools, and interact with people and other systems. My research asks how linguistic insight can help make these systems trustworthy across languages and contexts.

A connected research programme

From linguistic foundations to trustworthy AI systems.

My research follows a single arc: from understanding language and meaning, through studying and improving trustworthy language technologies, to extending these principles to increasingly complex AI systems.

  1. Linguistic foundations

    Understand how models represent meaning and linguistic structure across languages and cultures.

  2. Trustworthy systems

    Study whether systems are secure, privacy-preserving, safe, and factually reliable—and develop evaluation methods that make those claims defensible across languages and contexts.

  3. Next frontier

    Extend these principles to AI systems that retrieve knowledge, use tools, maintain context, and interact through language.

Established research

Four established research strands anchor the programme.

Two strands focus directly on core trustworthiness questions: language-model security and privacy, and factuality and evaluation. Multilingual NLP and formal and computational semantics provide foundations that run across the programme. Across all four, evaluation tests whether claims hold across languages and contexts.

Research strand

Language-model security and privacy

Linguistic perspectives on language-model security and privacy.

I study security and privacy risks in language models and text embeddings, including inversion attacks, information leakage, poisoning, and manipulated or machine-generated text. A central question is how these risks and possible defences vary across languages.

Information leakage

When can embeddings or model outputs be inverted to recover sensitive source text, and how do these risks vary across languages?

Poisoning and manipulation

How do linguistic properties interact with poisoning, model confusion, and the detection of manipulated or machine-generated text?

Selected papersChen et al., ACL 2024Lent et al., TACL 2025

Research strand

Factuality and evaluation

Reliable AI requires factual grounding and defensible evaluation.

I study how knowledge grounding, data quality, benchmark design, and language sampling shape claims about output factuality and model generalisation. This includes evaluating hallucinations, auditing multilingual data, and designing benchmarks whose conclusions hold beyond narrow sets of languages and settings.

Factual grounding

How can structured knowledge and provenance help detect, evaluate, and reduce unsupported model outputs?

Defensible evaluation

How should datasets, language sampling, and metrics be designed so that claims about model reliability generalise beyond narrow benchmarks?

Selected papersLavrinovics et al., JWS 2025Ploeger et al., CL 2026

Research strand

Multilingual NLP

Multilingual does not automatically mean inclusive.

Multilingual NLP remains concentrated in a relatively small set of high-resource languages. My work uses linguistic typology, careful language sampling, and targeted adaptation to understand cross-lingual generalisation and improve outcomes across a broader range of languages.

Typology and generalisation

How can typological evidence guide language sampling and cross-lingual transfer—and reveal when multilingual claims do not generalise?

Low-resource languages

How can data quality, targeted adaptation, and linguistic knowledge improve outcomes where digital resources are scarce?

Selected papersBjerva et al., CL 2019Lent et al., TACL 2024

Research strand

Formal and computational semantics

Linguistic structure can make model behaviour easier to analyse.

My work draws on formal and computational semantics to investigate how models encode meaning and context, where inference fails, and how explicit linguistic structure can support evaluation, interpretability, and safety analysis.

Meaning and inference

What aspects of meaning, inference, and context are captured by current language models, and where do their representations remain inadequate?

Linguistic analysis

How can syntax, morphology, semantics, and discourse help explain model behaviour beyond aggregate performance scores?

Selected papersAbzianidze et al., EACL 2017Bjerva et al., SemEval 2014

Funded research

Projects and programmes

Selected externally and institutionally funded research, the teams it supports, and work linked to each programme. VBN snapshots are current to July 2026.

Portfolio represented here

≈ DKK 35M

PI awards represented

2022–2026

Award years represented

8

Awards represented

Rounded across seven programmes and eight awards. The total includes internal AAU and co-led funding; the Google Award for Inclusion Research is shown with the Carlsberg programme. Project periods extend to 2030.

Funding

DKK 9.9M

Data Science Investigator (Ascending) · NNF Research Leader Programme

Principal Investigator

security2025–2030

(LM)²-SEC

Linguistically Motivated Language Model Security

Developing linguistically grounded approaches to security and privacy in language models, including inversion attacks, information leakage, and malicious manipulation across languages.

Funder
Novo Nordisk Foundation
Funded staffing
2 PhDs · 72 PD-months

Linked work

4 linked outputs

The first outputs address personal-information memorisation, extraction from diffusion language models, and semantic leakage from image embeddings, including work appearing at ACL 2026.

Project and linked outputs on VBN

Funding

DKK 6.2M

Sapere Aude: DFF-Research Leader

Principal Investigator

security2026–2030

TRUST

Building TRUST in Text: Linguistically Motivated Language Model Detection

Investigating whether linguistic signals in generated text can reveal poisoning and support more robust language models.

Funder
Independent Research Fund Denmark · Sapere Aude
Funded staffing
2 PhDs · 12 PD-months

Current status

Started May 2026

The four-year programme began in May 2026. Project-linked research outputs are forthcoming.

Project details on VBN

Funding

DKK 2.4M

Research Grant · Coefficient Giving

Principal Investigator

semantics2026–2028

Formal Semantic Methods for AI Safety

Formal and computational semantics for AI safety

Developing formal and computational semantic methods to analyse and explain language-model behaviour in multilingual settings.

Funder
Coefficient Giving
Funded staffing
36 PD-months

Current status

Started May 2026

The two-year programme began in May 2026. Project-linked research outputs are forthcoming.

Project details on VBN

Funding

~DKK 7M

AI:X Labs · Aalborg University

Lab Director / Principal Investigator · Co-led with Qiongxiu Li

security2025–2029

AI:SECURITY

AI-enabled threats and secure AI systems

Combining NLP and cybersecurity research to study AI-enabled threats and the security of language models in societal applications.

Funder
AI:X Labs · Aalborg University
Funded staffing
4 PhDs · main supervisor for 2

Linked work

5 linked outputs

Linked work spans multilingual memorisation, personal-information leakage, privacy-preserving graph aggregation, and safer model sharing, with papers at EMNLP 2025, ACL 2026, and ICASSP 2026.

Project and linked outputs on VBN

Funding

DKK 1.1M

Industrial PhD · Industrial Researcher Programme

Principal Investigator

factuality2024–2027

Guarantees of Factuality in LLM-based Extraction of Financial KPIs

Industrial PhD with Alipes Capital

Research on factual, knowledge-grounded extraction of financial indicators from earnings transcripts and company filings.

Funder
Innovation Fund Denmark
Funded staffing
1 PhD

Linked work

1 dataset · 1 paper

The project has delivered HiFi-KPI: a public dataset and accompanying paper for hierarchical KPI extraction from financial filings, linking applied research to reusable infrastructure.

Project and linked outputs on VBN

Funding

DKK 5M

Semper Ardens: Accelerate · Carlsberg Foundation

Principal Investigator

multilingual2022–2026

Multilingual Modelling for Resource-Poor Languages

Semper Ardens: Accelerate

Fundamental research on linguistic typology, multilingual modelling, and evaluation for languages underserved by current NLP systems.

Funder
Carlsberg Foundation
Funded staffing
3 PhDs · 36 PD-months

Linked work

40 outputs · 35 activities

VBN currently links 40 outputs spanning typological diversity, multilingual embedding inversion, and low-resource evaluation, including publications at ACL, EMNLP, TACL, and ICML.

Project and linked outputs on VBN

Funding

DKK 3M

Villum Synergy · Villum Foundation

Principal Investigator / NLP methods lead

education2024–2026

Digital Twins for Abundant Feedback

Novel Feedback Paradigms via Explainable Multilingual NLP

Studying explainable multilingual NLP for scalable, high-quality feedback in education.

Funder
Villum Foundation · Synergy
Funded staffing
24 PD-months

Linked work

13 outputs · 7 media items

Linked work connects synthetic educational feedback with LLM agents and knowledge-grounded factuality, alongside sustained public discussion of AI-supported teaching and assessment.

Project and linked outputs on VBN

Staffing figures describe funded positions; “PD-months” means funded postdoctoral researcher months. VBN totals are a dated snapshot and will grow as new outputs are linked.