Preprint
Large Language Models

Knowledge Engineering Using Large Language Models

Freitas, Tiago Carvalho(University of Minho), Costa Neto, Alvaro(Federal Institute of São Paulo), Pereira, Maria João Varanda(Robotics Research (United States)), Henriques, Pedro Rangel(University of Minho)
January 1, 2023DROPS (Schloss Dagstuhl – Leibniz Center for Informatics)269 citations

269

Citations

1

Influential Citations

DROPS (Schloss Dagstuhl – Leibniz Center for Informatics)

Venue

2023

Year

Abstract

Knowledge engineering is a discipline that focuses on the creation and maintenance of processes that generate and apply knowledge. Traditionally, knowledge engineering approaches have focused on knowledge expressed in formal languages. The emergence of large language models and their capabilities to effectively work with natural language, in its broadest sense, raises questions about the foundations and practice of knowledge engineering. Here, we outline the potential role of LLMs in knowledge engineering, identifying two central directions: 1) creating hybrid neuro-symbolic knowledge systems; and 2) enabling knowledge engineering in natural language. Additionally, we formulate key open research questions to tackle these directions.

Analysis

Why This Paper Matters

This paper addresses a critical juncture in knowledge engineering, where traditional formal language-based approaches are being challenged by the rise of large language models (LLMs). By outlining how LLMs can enable hybrid neuro-symbolic systems and natural language knowledge engineering, the authors provide a roadmap for integrating modern AI capabilities into a field that has long relied on structured representations. The paper's significance lies in its timely recognition that LLMs can democratize knowledge engineering, making it accessible to non-experts through natural language interfaces while also enhancing the expressiveness and adaptability of knowledge systems.

The paper also highlights the tension between the precision of formal languages and the flexibility of natural language, proposing that hybrid systems can combine the best of both worlds. This is particularly important for applications such as enterprise knowledge management, scientific discovery, and decision support systems, where both accuracy and ease of use are paramount. By formulating open research questions, the paper sets a clear agenda for future work, encouraging the community to explore these directions systematically.

Technical Contributions

The paper's primary technical contributions are conceptual rather than empirical, focusing on two key directions:

  • Hybrid neuro-symbolic knowledge systems: Combining neural LLMs with symbolic reasoning to leverage the strengths of both paradigms, such as using LLMs for natural language understanding and symbolic systems for logical inference.
  • Natural language knowledge engineering: Enabling the creation, maintenance, and application of knowledge directly in natural language, reducing the need for formal ontologies or rule-based representations.
  • Open research questions: The paper identifies several key questions, including how to ensure consistency and correctness in natural language knowledge bases, how to integrate LLMs with existing symbolic tools, and how to evaluate the quality of LLM-generated knowledge.

Results

As a position paper, this work does not present experimental results or quantitative metrics. Instead, it offers a conceptual framework and a set of research directions. The paper's impact is measured by its citation count (269) and its role in shaping subsequent research in the intersection of LLMs and knowledge engineering.

Significance

This paper has broader implications for the AI field by bridging the gap between natural language processing and knowledge representation. It challenges the traditional view that knowledge must be expressed in formal languages, opening up new possibilities for more intuitive and scalable knowledge systems. The proposed hybrid neuro-symbolic approach could lead to more robust AI systems that combine the flexibility of LLMs with the reliability of symbolic reasoning. Additionally, the focus on natural language knowledge engineering could lower the barrier to entry for domain experts who are not trained in formal knowledge representation, potentially accelerating knowledge-driven innovation across various fields.