Dr Yulong Chen

Dr Yulong Chen
Dr Yulong Chen
Dr Yulong Chen

Lecturer

Accepting PhDs

About

Biography

Yulong Chen is a Lecturer in the School of Natural and Computing Sciences at the University of Aberdeen. Before joining Aberdeen, he was an Affiliated Lecturer and Research Associate in the Department of Computer Science and Technology at the University of Cambridge, where he worked with Prof Andreas Vlachos. He received his PhD from Zhejiang University and Westlake University, advised by Prof Yue Zhang, and his MSc from the University of Edinburgh, advised by Prof Bonnie Webber.

His research focuses on natural language processing (NLP) and large language models (LLMs), with particular interests in model evaluation, automated fact-checking, text generation, and multilingual and low-resource NLP. He develops datasets, benchmarks, diagnostic frameworks, and modelling methods to better understand the capabilities and limitations of language technologies.

He has published more than 30 research papers, and his work has received over 3,000 citations. His research has involved collaborations with academic institutions, industrial research teams, and fact-checking organisations, including the University of Cambridge, Yale University, the University of Michigan, the University of Oxford, Microsoft, and members of the FEVER and AVeriTeC communities.

He is an Associate Editor of the Journal of Large Language Models and a reviewer for leading journals, such as Transactions of the Association for Computational Linguistics (TACL). He has also served as a Senior Action Editor for ACL Rolling Review (ARR), and a Senior Area Chair and Area Chair for top-tier NLP and AI conferences such as ACL and EMNLP. In addition, he served on the organising committee of NLPCC 2026 and has co-organised workshops and shared tasks, including the FEVER workshops and the DialogSum Challenge.

External Memberships

Visiting Researcher, Department of Computer Science and Technology, University of Cambridge

Research

Research Overview

My research lies at the intersection of natural language processing (NLP), large language models (LLMs), and trustworthy artificial intelligence (AI). I am particularly interested in understanding what language models can and cannot do, how their capabilities can be evaluated reliably, and how they can be developed for applications requiring accurate, verifiable, and transparent outputs.

My research covers LLM evaluation, automated fact-checking, factuality and hallucination, summarisation, multilingual and low-resource NLP, and language model generalisation. A recurring theme across my work is the development of datasets, benchmarks, diagnostic methods, and evaluation frameworks that reveal model limitations which may be hidden by conventional aggregate performance measures.

I am also interested in translating fundamental research into practical systems, particularly for fact-checking and information verification.

Research Areas

Accepting PhDs

I am currently accepting PhDs in Computing Science.

Please get in touch if you would like to discuss your research ideas further.

Computing Science

  • Supervising
  • Accepting PhDs

Research Specialisms

  • Natural Language Processing
  • Artificial Intelligence

Our research specialisms are based on the Higher Education Classification of Subjects (HECoS) which is HESA open data, published under the Creative Commons Attribution 4.0 International licence.

Current Research

My current research investigates the evaluation, reliability, and generalisation of LLMs.

I work on methods for identifying model weaknesses that may not be captured by conventional benchmarks, including weaknesses in reasoning, language understanding, factual accuracy, fairness, and cross-lingual generalisation.

I also study automated fact-checking and information verification, including evidence retrieval, claim verification, factuality evaluation, and systems that support the analysis of information across multiple sources.

More broadly, I am interested in trustworthy and responsible language technologies, particularly the design of evaluation methods that reflect how models behave in real-world settings.

Past Research

My previous research focused primarily on dialogue and document summarisation, multilingual NLP, and learning in low-resource settings.

This work included the construction of datasets and shared tasks, the development of few-shot and cross-lingual summarisation methods, and the evaluation of generated text.

I have also worked on prompt-based learning, information extraction, semantic representation, parsing, machine translation, and other areas of NLP.

Knowledge Exchange

I contribute to the international research community through the organisation of workshops, shared tasks, and evaluation campaigns. I have co-organised the FEVER workshops at EMNLP 2024, ACL 2025, and EACL 2026, and I led the organisation of the DialogSum Challenge at INLG 2022. These initiatives bring together academic researchers, industry practitioners, and fact-checking organisations to develop shared datasets, evaluation standards, and reproducible research practices.

I also contribute to research quality and community development through editorial and conference leadership. My roles have included Associate Editor for the Journal of Large Language Models, Senior Action Editor for ACL Rolling Review, and Area Chair or Senior Area Chair for major international conferences in NLP and AI.

I am particularly interested in knowledge-exchange activities with fact-checking organisations, technology companies, media organisations, educators, and public-sector bodies seeking to evaluate or deploy language technologies responsibly.

Collaborations

My academic collaborations have involved researchers at the University of Cambridge, the University of Michigan, the University of Oxford, the University of Edinburgh, Peking University, Zhejiang University, Westlake University, etc.

I have also collaborated with industrial researchers from Microsoft, Tencent, OpenAI, etc. During my work with Microsoft Azure Cognitive Services Research, I contributed to research on few-shot text generation, prompt-based learning, and model evaluation, with research outcomes contributing to product development. I have also worked with Language Bridge on activities connecting language and AI research with machine translation and multilingual NLP. In automated fact-checking, I have collaborated with researchers and practitioners from Full Fact and the FEVER and AVeriTeC communities.

I welcome collaborations in LLM evaluation, trustworthy AI, automated fact-checking, misinformation, factuality, multilingual NLP, summarisation, retrieval-augmented generation, and human-centred evaluation.