LangChain学习之 Question And Answer的操作

2024-06-04 23:04

本文主要是介绍LangChain学习之 Question And Answer的操作,希望对大家解决编程问题提供一定的参考价值,需要的开发者们随着小编来一起学习吧!

1. 学习背景

在LangChain for LLM应用程序开发中课程中,学习了LangChain框架扩展应用程序开发中语言模型的用例和功能的基本技能,遂做整理为后面的应用做准备。视频地址:基于LangChain的大语言模型应用开发+构建和评估。

2. Q&A的作用

基于文档的问答系统是LLM的典型应用,给定一段可能从PDF文件、网页或某公司的内部文档库中提取的文本,可以使用LLM检索文档对问题进行回答。以下代码基于jupyternotebook运行。

1.导入环境

import osfrom dotenv import load_dotenv, find_dotenv
_ = load_dotenv(find_dotenv()) # read local .env file
from langchain.chains import RetrievalQA
from langchain.chat_models import ChatOpenAI
from langchain.document_loaders import CSVLoader
from langchain.vectorstores import DocArrayInMemorySearch
from IPython.display import display, Markdown

2.2 读取数据进行查询

from langchain.indexes import VectorstoreIndexCreator
# 没有docarray环境需要安装。命令:!pip install docarray# 要用到的数据文件
file = 'OutdoorClothingCatalog_1000.csv'
loader = CSVLoader(file_path=file, encoding='utf-8')# 此处我们已完成了文档的向量存储
index = VectorstoreIndexCreator(vectorstore_cls=DocArrayInMemorySearch).from_loaders([loader])# 创建提问语句
query ="Please list all your shirts with sun protection in a table in markdown and summarize each one."# 传入query内容,使用index生成响应
response = index.query(query)# 以markdown方式进行呈现,注意LLM生成的样式可能存在差异
display(Markdown(response))

输出如下:

NameDescriptionSun Protection Rating
Men’s Tropical Plaid Short-Sleeve ShirtMade of 100% polyester, UPF 50+ rating, front and back cape venting, two front bellows pocketsSPF 50+, blocks 98% of harmful UV rays
Men’s Plaid Tropic Shirt, Short-SleeveMade of 52% polyester and 48% nylon, UPF 50+ rating, front and back cape venting, two front bellows pocketsSPF 50+, blocks 98% of harmful UV rays
Men’s TropicVibe Shirt, Short-SleeveMade of 71% nylon and 29% polyester, UPF 50+ rating, front and back cape venting, two front bellows pocketsSPF 50+, blocks 98% of harmful UV rays
Sun Shield ShirtMade of 78% nylon and 22% Lycra Xtra Life fiber, UPF 50+ rating, wicks moisture, abrasion resistantSPF 50+, blocks 98% of harmful UV rays

All four shirts provide UPF 50+ sun protection, blocking 98% of the sun’s harmful rays. The Men’s Tropical Plaid Short-Sleeve Shirt is made of 100% polyester and is wrinkle-resistant。

至此,内容已经查出来了,并生成了一小段总结的话。那么底层的原理又是什么呢?

2.3 底层原理

2.3.1向量化

一般的大模型一次只能接收几千个单词,如图:
在这里插入图片描述
如果有个很大的文档,我们要怎样让LLM对文档进行问答呢?这里就需要Embedding和向量存储发挥作用了。
在这里插入图片描述
什么是Embedding?Embedding将一段文本转换成数字,用一组数字表示这段文本。这组数字捕捉了它所代表的文字片段的肉容含义。内容相似的文本片段会有相似的向量值,这样我们可以在向量空间中比较文本片段。例如,我们有三段话:

  1. My dog Rover likes to chase squirrels.
  2. Fluffy, my cat, refuses to eat from a can.
  3. The Chevy Bolt accelerates to 60 mph in 6.7 seconds.

三段话前两个描述宠物,第三个描述汽车,向量化后如图:
在这里插入图片描述
如果我们观察数值空间中的表示,可以看到当我们比较关于两个关于宠物的句子的向量时,它们相似度非常高。将其与汽车相关的语句进行比对,可以看到相关程度非常低。利用向量可以很轻松的让我们找出哪些片段是相似的。利用这种技术,我们可以从文档中找出与提问相似的片段,传递给LLM进行解答。

2.3.2向量数据库

在这里插入图片描述
向量数据库是一种存储方法,可以存储我们在前面创建的那种矢量数字数组。往向量数据库中新建数据的方式,就是将文档拆分成块,每块生成Embedding,然后把Embedding和原始块一起存储到数据库中。

因为有些大文档无法整个传给文档,因此要先切块,然后只把最相关的内容存入,然后,把每个文本块生成一个Embedding,然后将这些Embedding存储在向量数据库中。如图:
在这里插入图片描述
当查询过来,我们先将查询内容embedding,得到一个数组,然后将这个数字数组与向量数据库中的所有向量进行比较,选择最相似的前若干个文本块。

拿到这些文本块后,将这些文本块和原始的查询内容一起传递给语言模型,这样可以让语言模型根据检索出来的文档内容生成最终答案。

2.4 再了解底层原理

loader = CSVLoader(file_path=file, encoding='utf-8')
docs = loader.load()
docs[0]

输出如下:

Document(page_content=": 0\nname: Women's Campside Oxfords\ndescription: This ultracomfortable lace-to-toe Oxford boasts a super-soft canvas, thick cushioning, and quality construction for a broken-in feel from the first time you put them on. \n\nSize & Fit: Order regular shoe size. For half sizes not offered, order up to next whole size. \n\nSpecs: Approx. weight: 1 lb.1 oz. per pair. \n\nConstruction: Soft canvas material for a broken-in feel and look. Comfortable EVA innersole with Cleansport NXT® antimicrobial odor control. Vintage hunt, fish and camping motif on innersole. Moderate arch contour of innersole. EVA foam midsole for cushioning and support. Chain-tread-inspired molded rubber outsole with modified chain-tread pattern. Imported. \n\nQuestions? Please contact us for any inquiries.", metadata={'source': 'OutdoorClothingCatalog_1000.csv', 'row': 0})

接着

# 使用OpenAIEmbeddings完成embedding
from langchain.embeddings import OpenAIEmbeddings
embeddings = OpenAIEmbeddings()
#使用embed_query模拟生成embeddings向量
embed = embeddings.embed_query("Hi my name is Harrison")
print(len(embed))
print(embed[:5])

输出如下:

1536[-0.021900920197367668, 0.006746490020304918, -0.018175246194005013, -0.039119575172662735, -0.014097143895924091]

可以看到,embedding向量的长度为1536,数组的前五个向量如上。

# 接着我们将刚刚加载的所有文本片段生成Embedding,并将它们存储在一个向量数据库中
db = DocArrayInMemorySearch.from_documents(docs, embeddings
)
# 创建对话查询语句
query = "Please suggest a shirt with sunblocking"
# 向量数据库中使用similarity_search方法得到查询的文档列表
docs = db.similarity_search(query)
print(len(docs))
print(docs[0])

输出如下:

4
Document(page_content=': 255\nname: Sun Shield Shirt by\ndescription: "Block the sun, not the fun – our high-performance sun shirt is guaranteed to protect from harmful UV rays. \n\nSize & Fit: Slightly Fitted: Softly shapes the body. Falls at hip.\n\nFabric & Care: 78% nylon, 22% Lycra Xtra Life fiber. UPF 50+ rated – the highest rated sun protection possible. Handwash, line dry.\n\nAdditional Features: Wicks moisture for quick-drying comfort. Fits comfortably over your favorite swimsuit. Abrasion resistant for season after season of wear. Imported.\n\nSun Protection That Won\'t Wear Off\nOur high-performance fabric provides SPF 50+ sun protection, blocking 98% of the sun\'s harmful rays. This fabric is recommended by The Skin Cancer Foundation as an effective UV protectant.', metadata={'source': 'OutdoorClothingCatalog_1000.csv', 'row': 255})

可以看到,得到了4个相关的文档列表内容,第一个内容如上所示。

2.5 如何利用这个来回答得到提问的结果

# 首先,需要从这个向量存储器创建一个检索器(Retriever)
retriever = db.as_retriever()
# 定义一个LLM模型
llm = ChatOpenAI(temperature = 0.0)
# 手动将检索出来的内容合并成一段话
qdocs = "".join([docs[i].page_content for i in range(len(docs))])
# 将提问和检索出来的内容一起交给LLM,并让其生成一段摘要
response = llm.call_as_llm(f"{qdocs} Question: Please list all your \
shirts with sun protection in a table in markdown and summarize each one.") 
display(Markdown(response))

输出如下:

NameDescription
Sun Shield ShirtHigh-performance sun shirt with UPF 50+ sun protection, moisture-wicking, and abrasion-resistant fabric. Fits comfortably over swimsuits. Recommended by The Skin Cancer Foundation.
Men’s Plaid Tropic ShirtUltracomfortable shirt with UPF 50+ sun protection, wrinkle-free fabric, and front/back cape venting. Made with 52% polyester and 48% nylon.
Men’s TropicVibe ShirtMen’s sun-protection shirt with built-in UPF 50+ and front/back cape venting. Made with 71% nylon and 29% polyester.
Men’s Tropical Plaid Short-Sleeve ShirtLightest hot-weather shirt with UPF 50+ sun protection, front/back cape venting, and two front bellows pockets. Made with 100% polyester and is wrinkle-resistant.

All of these shirts provide UPF 50+ sun protection, blocking 98% of the sun’s harmful rays. They are made with high-performance fabrics that are moisture-wicking, abrasion-resistant, and/or wrinkle-free. Some have front/back cape venting for added comfort in hot weather. The Sun Shield Shirt is recommended by The Skin Cancer Foundation.

2.6使用langchain进行封装运行

qa_stuff = RetrievalQA.from_chain_type(llm=llm, chain_type="stuff", retriever=retriever, verbose=True
)
query =  "Please list all your shirts with sun protection in a table in markdown and summarize each one."
response = qa_stuff.run(query)

输出如下:

Shirt NameDescription
Men’s Tropical Plaid Short-Sleeve ShirtRated UPF 50+ for superior protection from the sun’s UV rays. Made of 100% polyester and is wrinkle-resistant. With front and back cape venting that lets in cool breezes and two front bellows pockets. Provides the highest rated sun protection possible.
Men’s Plaid Tropic Shirt, Short-SleeveRated to UPF 50+, helping you stay cool and dry. Made with 52% polyester and 48% nylon, this shirt is machine washable and dryable. Additional features include front and back cape venting, two front bellows pockets and an imported design. With UPF 50+ coverage, you can limit sun exposure and feel secure with the highest rated sun protection available.
Men’s TropicVibe Shirt, Short-SleeveBuilt-in UPF 50+ has the lightweight feel you want and the coverage you need when the air is hot and the UV rays are strong. Made with Shell: 71% Nylon, 29% Polyester. Lining: 100% Polyester knit mesh. Wrinkle resistant. Front and back cape venting lets in cool breezes. Two front bellows pockets. Imported.
Sun Shield ShirtHigh-performance sun shirt is guaranteed to protect from harmful UV rays. Made with 78% nylon, 22% Lycra Xtra Life fiber. Fits comfortably over your favorite swimsuit. Abrasion resistant for season after season of wear.

All of the shirts listed have sun protection with a UPF rating of 50+ and block 98% of the sun’s harmful rays. The Men’s Tropical Plaid Short-Sleeve Shirt is made of 100% polyester and has front and back cape venting and two front bellows pockets. The Men’s Plaid Tropic Shirt, Short-Sleeve is made with 52% polyester and 48% nylon and has front and back cape venting and two front bellows pockets. The Men’s TropicVibe Shirt, Short-Sleeve is made with Shell: 71% Nylon, 29% Polyester. Lining: 100% Polyester knit mesh and has front and back cape venting and two front bellows pockets. The Sun Shield Shirt is made with 78% nylon, 22% Lycra Xtra Life fiber and fits comfortably over your favorite swimsuit.

同样的,我们尝试用index.query也会得到同样的内容。

response = index.query(query, llm=llm)

输出结果和之前的一致

3.总结

Q&A可以用一行代码完成,也可以把它分成五个详细的步骤,可以查看每一步的详细结果。五个步骤可以详细的让我们理解到它底层到底是如何执行的。此外,chain_type="stuff" 参数还有其他三种,可以根据实际情况选取合适的参数,另外三种如图,有需要可以根据实际情况选取合适的参数进行实验。
在这里插入图片描述

这篇关于LangChain学习之 Question And Answer的操作的文章就介绍到这儿,希望我们推荐的文章对编程师们有所帮助!



http://www.chinasem.cn/article/1031372

相关文章

SQL中JOIN操作的条件使用总结与实践

《SQL中JOIN操作的条件使用总结与实践》在SQL查询中,JOIN操作是多表关联的核心工具,本文将从原理,场景和最佳实践三个方面总结JOIN条件的使用规则,希望可以帮助开发者精准控制查询逻辑... 目录一、ON与WHERE的本质区别二、场景化条件使用规则三、最佳实践建议1.优先使用ON条件2.WHERE用

Linux链表操作方式

《Linux链表操作方式》:本文主要介绍Linux链表操作方式,具有很好的参考价值,希望对大家有所帮助,如有错误或未考虑完全的地方,望不吝赐教... 目录一、链表基础概念与内核链表优势二、内核链表结构与宏解析三、内核链表的优点四、用户态链表示例五、双向循环链表在内核中的实现优势六、典型应用场景七、调试技巧与

Go学习记录之runtime包深入解析

《Go学习记录之runtime包深入解析》Go语言runtime包管理运行时环境,涵盖goroutine调度、内存分配、垃圾回收、类型信息等核心功能,:本文主要介绍Go学习记录之runtime包的... 目录前言:一、runtime包内容学习1、作用:① Goroutine和并发控制:② 垃圾回收:③ 栈和

Java Multimap实现类与操作的具体示例

《JavaMultimap实现类与操作的具体示例》Multimap出现在Google的Guava库中,它为Java提供了更加灵活的集合操作,:本文主要介绍JavaMultimap实现类与操作的... 目录一、Multimap 概述Multimap 主要特点:二、Multimap 实现类1. ListMult

Python中文件读取操作漏洞深度解析与防护指南

《Python中文件读取操作漏洞深度解析与防护指南》在Web应用开发中,文件操作是最基础也最危险的功能之一,这篇文章将全面剖析Python环境中常见的文件读取漏洞类型,成因及防护方案,感兴趣的小伙伴可... 目录引言一、静态资源处理中的路径穿越漏洞1.1 典型漏洞场景1.2 os.path.join()的陷

Android学习总结之Java和kotlin区别超详细分析

《Android学习总结之Java和kotlin区别超详细分析》Java和Kotlin都是用于Android开发的编程语言,它们各自具有独特的特点和优势,:本文主要介绍Android学习总结之Ja... 目录一、空安全机制真题 1:Kotlin 如何解决 Java 的 NullPointerExceptio

Python使用Code2flow将代码转化为流程图的操作教程

《Python使用Code2flow将代码转化为流程图的操作教程》Code2flow是一款开源工具,能够将代码自动转换为流程图,该工具对于代码审查、调试和理解大型代码库非常有用,在这篇博客中,我们将深... 目录引言1nVflRA、为什么选择 Code2flow?2、安装 Code2flow3、基本功能演示

Python中OpenCV与Matplotlib的图像操作入门指南

《Python中OpenCV与Matplotlib的图像操作入门指南》:本文主要介绍Python中OpenCV与Matplotlib的图像操作指南,本文通过实例代码给大家介绍的非常详细,对大家的学... 目录一、环境准备二、图像的基本操作1. 图像读取、显示与保存 使用OpenCV操作2. 像素级操作3.

python操作redis基础

《python操作redis基础》Redis(RemoteDictionaryServer)是一个开源的、基于内存的键值对(Key-Value)存储系统,它通常用作数据库、缓存和消息代理,这篇文章... 目录1. Redis 简介2. 前提条件3. 安装 python Redis 客户端库4. 连接到 Re

Java Stream.reduce()方法操作实际案例讲解

《JavaStream.reduce()方法操作实际案例讲解》reduce是JavaStreamAPI中的一个核心操作,用于将流中的元素组合起来产生单个结果,:本文主要介绍JavaStream.... 目录一、reduce的基本概念1. 什么是reduce操作2. reduce方法的三种形式二、reduce