Domain-specific chatbots for science using embeddings

Kevin G. Yager

doi:10.1039/D3DD00112A

You do not have JavaScript enabled. Please enable JavaScript to access the full features of the site or access our non-JavaScript page.

Domain-specific chatbots for science using embeddings†

Kevin G. Yager

*^a

Author affiliations

* Corresponding authors

^a Center for Functional Nanomaterials, Brookhaven National Laboratory, Upton, New York 11973, USA
E-mail: kyager@bnl.gov

Abstract

Large language models (LLMs) have emerged as powerful machine-learning systems capable of handling a myriad of tasks. Tuned versions of these systems have been turned into chatbots that can respond to user queries on a vast diversity of topics, providing informative and creative replies. However, their application to physical science research remains limited owing to their incomplete knowledge in these areas, contrasted with the needs of rigor and sourcing in science domains. Here, we demonstrate how existing methods and software tools can be easily combined to yield a domain-specific chatbot. The system ingests scientific documents in existing formats, and uses text embedding lookup to provide the LLM with domain-specific contextual information when composing its reply. We similarly demonstrate that existing image embedding methods can be used for search and retrieval across publication figures. These results confirm that LLMs are already suitable for use by physical scientists in accelerating their research efforts.

Download options Please wait...

Supplementary files

Supplementary information PDF (15021K)

Article information

DOI: https://doi.org/10.1039/D3DD00112A
Article type: Paper
Submitted: 12 Jun 2023
Accepted: 04 Oct 2023
First published: 10 Oct 2023
This article is Open Access

Download Citation

Digital Discovery, 2023,2, 1850-1861

Permissions

Request permissions

Domain-specific chatbots for science using embeddings

K. G. Yager, Digital Discovery, 2023, 2, 1850 DOI: 10.1039/D3DD00112A

This article is licensed under a Creative Commons Attribution-NonCommercial 3.0 Unported Licence. You can use material from this article in other publications, without requesting further permission from the RSC, provided that the correct acknowledgement is given and it is not used for commercial purposes.

To request permission to reproduce material from this article in a commercial publication, please go to the Copyright Clearance Center request page.

If you are an author contributing to an RSC publication, you do not need to request permission provided correct acknowledgement is given.

If you are the author of this article, you do not need to request permission to reproduce figures and diagrams provided correct acknowledgement is given. If you want to reproduce the whole article in a third-party commercial publication (excluding your thesis/dissertation for which permission is not required) please go to the Copyright Clearance Center request page.

Social activity

Fetching data from CrossRef.
This may take some time to load.

Digital Discovery

Domain-specific chatbots for science using embeddings†

Abstract

Supplementary files

Article information

Download Citation

Permissions

Domain-specific chatbots for science using embeddings

Social activity

Search articles by author

Spotlight

Advertisements