體與關(guān)系抽取模塊解析:從 Pydantic 模型到 LLM 提示詞組裝的全鏈路實(shí)踐)
人工智能AI AgentAgent 框架大模型工具調(diào)用RAG提示工程強(qiáng)化學(xué)習(xí)【免費(fèi)下載鏈接】agent-coreopenJiuwen agent-core可提供AI Agent開發(fā)、運(yùn)行、調(diào)優(yōu)與演進(jìn)相關(guān)的全套SDK能力項(xiàng)目地址https://gitcode.com/openJiuwen/agent-core點(diǎn)擊查看免費(fèi)下載本文圍繞 openJiuwen agent-core 圖記憶Graph Memory能力中的核心模塊openjiuwen.core.memory.graph.extraction展開它負(fù)責(zé)從對(duì)話、文檔與 JSON 三類來源中抽取實(shí)體與關(guān)系并為記憶圖提供結(jié)構(gòu)化輸入。讀完本文你將掌握該模塊的分層結(jié)構(gòu)多語言響應(yīng)模型基類MultilingualBaseModel、實(shí)體/關(guān)系類型定義、抽取用 Pydantic 輸出模型、線程安全的提示詞模板管理器、按情節(jié)類型組裝提示詞的入口函數(shù)以及 LLM 回復(fù)的 JSON 容錯(cuò)解析工具并能結(jié)合倉庫中的調(diào)用鏈與測試用例理解其在GraphMemory寫入流水線中的實(shí)際位置。一、模塊定位圖記憶流水線中的結(jié)構(gòu)化輸入引擎在 openJiuwen agent-core 中圖記憶由 graph_memory/base.py 中的GraphMemory類承載它維護(hù)一張覆蓋用戶對(duì)話與文檔的知識(shí)圖譜通過 LLM 抽取實(shí)體與關(guān)系、執(zhí)行合并與去重并支持實(shí)體/關(guān)系/情節(jié)的語義檢索。而extraction模塊正是這條流水線的輸入端——所有進(jìn)入記憶圖的實(shí)體與關(guān)系都先由它完成從原始文本到結(jié)構(gòu)化聲明的轉(zhuǎn)換。從調(diào)用關(guān)系看GraphMemory.add_memory的完整流程見 graph_memory/base.py依次經(jīng)過時(shí)區(qū)預(yù)測extract_timezone——為后續(xù)關(guān)系抽取中的相對(duì)時(shí)間解析提供時(shí)區(qū)參考實(shí)體聲明抽取extract_entity_declaration——識(shí)別當(dāng)前消息/文檔/JSON 中的實(shí)體名稱與類型關(guān)系抽取extract_relation_declaration——在實(shí)體列表之上抽取帶時(shí)間信息的事實(shí)三元組實(shí)體去重與合并dedupe_entity_list、merge_existing_entities——與庫中已有實(shí)體比對(duì)并歸并實(shí)體摘要與屬性抽取extract_entity_attributes關(guān)系過濾與去重filter_relations_for_merge、dedupe_relation_list。每一步的 LLM 調(diào)用都遵循同一范式由extraction_prompts中的組裝函數(shù)返回(模板變量, PromptTemplate, LLM 響應(yīng)格式)三元組交由_invoke_llm執(zhí)行最后用parse_response.parse_json將回復(fù)解析回結(jié)構(gòu)化數(shù)據(jù)見 graph_memory/base.py 中parse_all_relations(ensure_list(parse_json(...)))的典型組合。因此理解extraction模塊是掌握?qǐng)D記憶寫入鏈路的前提。二、MultilingualBaseModel多語言響應(yīng)模型基類openjiuwen.core.memory.graph.extraction.base.MultilingualBaseModel是所有抽取輸出模型的基類base.py。它基于 PydanticBaseModel在保留字段校驗(yàn)?zāi)芰Φ耐瑫r(shí)解決了兩個(gè)問題多語言描述替換與將輸出模型轉(zhuǎn)為 LLM 可消費(fèi)的字符串/結(jié)構(gòu)化格式。其多語言描述來源于模塊級(jí)字典MULTILINGUAL_DESCRIPTION由 prompts/entity_extraction/cn.py 與 prompts/entity_extraction/en.py 在調(diào)用register_language()時(shí)填充。模型字段中的description{{[rel_name]}}這類占位符會(huì)被替換為對(duì)應(yīng)語言的實(shí)際描述例如中文該實(shí)體聯(lián)系的名稱、英文Name of factual relation。2.1 multilingual_model_json_schema按語言生成 JSON Schemadef multilingual_model_json_schema(cls, language: str cn, strict: bool False, **kwargs) - dict[str, Any]languagestr可選語言標(biāo)識(shí)如cn、en默認(rèn)cnstrictbool可選是否啟用嚴(yán)格模式默認(rèn)Falsekwargs透傳給 Pydanticmodel_json_schema的額外參數(shù)。實(shí)現(xiàn)上base.py先生成標(biāo)準(zhǔn) Pydantic schema再用_recursive_replace把description按MULTILINGUAL_DESCRIPTION[language]遞歸替換。strictTrue時(shí)通過 BFS 遍歷 schema 樹為所有type object且含properties的節(jié)點(diǎn)強(qiáng)制寫入additionalProperties: False并自動(dòng)補(bǔ)齊required字段列表——這正是 OpenAI 結(jié)構(gòu)化輸出格式的要求。對(duì)應(yīng)單元測試 tests/unit_tests/core/memory/graph/extraction/test_base.py 中_object_nodes_with_wrong_additional_properties專門校驗(yàn)了這一點(diǎn)。2.2 readable_schema生成 LLM 可讀的 schema 文本def readable_schema(cls, language: str cn, **kwargs) - tuple[str, dict]返回(schema 字符串, $defs 中 ref 的 properties 字典)。實(shí)現(xiàn)細(xì)節(jié)base.py基于multilingual_model_json_schema生成后移除title、required等冗余信息若存在$defs將$ref引用映射為類型名并單獨(dú)抽出 ref 的 properties 字典遍歷model_fields將每個(gè)字段渲染為字段名: JSON類型 # 多語言描述的形式。文檔中給出的示例輸出Fact.readable_schema(languagecn)驗(yàn)證了這一點(diǎn)其中valid_since被渲染為帶 ISO 格式提示的描述from openjiuwen.core.memory.graph.extraction.prompts import entity_extraction # 注冊(cè) cn/en from openjiuwen.core.memory.graph.extraction.extraction_models import Fact, RelationExtraction out_str, ref_dict Fact.readable_schema(languagecn) # name: str # 該實(shí)體聯(lián)系的名稱 # fact: str # 關(guān)于實(shí)體聯(lián)系的事實(shí) # valid_since: str # 事實(shí)/關(guān)系的生效日期請(qǐng)使用ISO格式Y(jié)YYY-MM-DDTHH:MM:SS[HH:MM] # valid_until: str # 事實(shí)/關(guān)系的中止日期請(qǐng)使用ISO格式Y(jié)YYY-MM-DDTHH:MM:SS[HH:MM] # source_id: int # 主體的實(shí)體ID # target_id: int # 客體的實(shí)體ID對(duì)于嵌套模型RelationExtraction.readable_schema(languageen)返回的字符串為extracted_relations: list[Fact] # List of extracted relations而 ref 字典的鍵為[Fact]其值為{name: {type: string, description: Name of factual relation}, ...}。生成的文本會(huì)被 prompts/entity_extraction/base.py 的format_schema_info以---\n# 輸出定義最終輸出需要為JSON\npython\n{out_str}\n的形式拼接到提示詞末尾讓 LLM 直接照著輸出格式生成 JSON。2.3 response_format轉(zhuǎn)換為 OpenAI 標(biāo)準(zhǔn)響應(yīng)格式def response_format(cls, language: str cn) - dict[str, Any]返回可直接作為 OpenAIresponse_format字段使用的字典base.py{ type: json_schema, json_schema: { schema: cls.multilingual_model_json_schema(language, strictTrue), name: cls.__name__, strict: False, }, }即schema 使用嚴(yán)格模式生成強(qiáng)制additionalProperties: False但響應(yīng)層strict保持False為推理類模型留出余地。每個(gè)組裝函數(shù)返回的三元組第三項(xiàng)即output_model.response_format(language)同時(shí)被用于 LLM 調(diào)用與后續(xù)parse_json的鍵過濾。三、實(shí)體與關(guān)系類型定義entity_type_definitionopenjiuwen.core.memory.graph.extraction.entity_type_definition提供抽取時(shí)的類型約束與展示所需的類型定義entity_type_definition.py。類基類默認(rèn)字段說明EntityDefAttrMultilingualBaseModelcontent: str 實(shí)體類型的屬性定義描述占位符{{[ent_summary]}}EntityDefBaseModelnameEntity、description多語言字典、attributes基礎(chǔ)實(shí)體類型定義RelationDefBaseModelnameRelation、description、lhs、rhs關(guān)系類型定義lhs/rhs為左右端實(shí)體類型HumanEntityEntityDefnameHuman表示用戶的實(shí)體類型AIEntityEntityDefnameAI表示AI 助手的實(shí)體類型各類型的多語言描述字典在 cn.py 中注冊(cè)ENTITY_DEFINITION_DESCRIPTION[cn] 默認(rèn)實(shí)體類型。若該實(shí)體不屬于其他提供的類型請(qǐng)選此類。HUMAN_ENTITY_DESCRIPTION[cn] 代表人類的實(shí)體類型可以是用戶也可以是其他人。AI_ENTITY_DESCRIPTION[cn] 代表AI的實(shí)體類型可能是聊天助手也可能是其他智能體。英文側(cè)對(duì)應(yīng)的描述為Default entity type, pick this if no other option is suitable.等見 en.py。RelationDef的中文描述模板為{name}{lhs}-[{name}]-{rhs}{description}見 base.py 的format_relation_definitions以清晰的(lhs)-[relation]-(rhs)形式約束關(guān)系兩端。在GraphMemory初始化時(shí)默認(rèn)實(shí)體類型集合為[EntityDef(), HumanEntity(), AIEntity()]EntityDeclaration.entity_type_id即對(duì)應(yīng)此列表的下標(biāo)。四、抽取輸出模型extraction_modelsopenjiuwen.core.memory.graph.extraction.extraction_models定義抽取各階段的結(jié)果模型extraction_models.py全部繼承MultilingualBaseModel4.1 基礎(chǔ)數(shù)據(jù)結(jié)構(gòu)模型字段用途Datetimeyear/month/day/hour/minute/secondint表示日期時(shí)間當(dāng)前未使用字段帶多語言 descriptionEntityDeclarationname: str、entity_type_id: int單條實(shí)體聲明類型 id 對(duì)應(yīng)EntityDef列表下標(biāo)Duplicationname: str、id: int、duplicate_ids: list[int]實(shí)體去重結(jié)果代表名、保留 id、重復(fù) id 列表Factname/fact/valid_since/valid_until: str、source_id/target_id: int一條事實(shí)關(guān)系源/目標(biāo)實(shí)體 id 為 1 基編號(hào)PossibleTimezonename: str、offset_from_utc: str、reasoning: str時(shí)區(qū)猜測名稱、UTC 偏移HH:MM、推理說明4.2 各階段輸出模型模型字段對(duì)應(yīng)階段EntityExtractionextracted_entities: list[EntityDeclaration]實(shí)體聲明抽取EntitySummarysummary: str、attributes: dict實(shí)體摘要與屬性抽取EntityDuplicationduplicated_entities: list[Duplication]實(shí)體去重RelationExtractionextracted_relations: list[Fact]關(guān)系抽取RelevantFactsbrief_reasoning: str、relevant_relations: list[int]關(guān)系過濾合并前篩選相關(guān)關(guān)系 idTimezonePredictionsextracted_relations: list[PossibleTimezone]時(shí)區(qū)預(yù)測字段名沿用關(guān)系以兼容提示MergeRelationsneed_merging: bool、short_reasoning: str、combined_content: str、duplicate_ids: list[int]、valid_since/valid_until: str關(guān)系合并字段描述全部使用{{[key]}}占位符由多語言注冊(cè)表在生成 schema 時(shí)替換。以Fact為例中文描述為rel_name→該實(shí)體聯(lián)系的名稱、rel_fact→關(guān)于實(shí)體聯(lián)系的事實(shí)、rel_valid_since/rel_valid_until→事實(shí)/關(guān)系的生效日期/中止日期請(qǐng)使用ISO格式Y(jié)YYY-MM-DDTHH:MM:SS[HH:MM]、rel_source_id/rel_target_id→主體的實(shí)體ID/客體的實(shí)體ID見 cn.py。值得注意的細(xì)節(jié)EntitySummary.attributes被建模為自由的dict而非固定 schema因此在嚴(yán)格模式下不會(huì)被強(qiáng)制additionalProperties: False見 test_base.py 的注釋說明為屬性鍵值對(duì)保留了靈活性。五、提示詞模板管理器與 .pr.md 格式5.1 ThreadSafePromptManageropenjiuwen.core.memory.graph.extraction.prompts.manager.ThreadSafePromptManager是線程安全的提示詞模板管理器manager.py以單例方式對(duì)外使用別名TemplateManager見 prompts/init.py。無參構(gòu)造首次初始化時(shí)通過glob.glob(..., recursiveTrue)掃描prompts/下所有**/*.pr.md文件按所在目錄批量注冊(cè)到內(nèi)部PromptMgrload_pr_content(content)staticmethod解析.pr.md原始內(nèi)容為消息列表。使用#user#、#system#、#assistant#、#tool#四種角色標(biāo)記正則PR_PATTERN每段內(nèi)容對(duì)應(yīng)一個(gè){role: ..., content: ...}get(name)按名稱獲取已注冊(cè)的PromptTemplate未注冊(cè)返回None例如模板名entity_extraction_conversation_cnregister_in_bulk(prompt_dir, name)將指定目錄下所有.pr.md注冊(cè)為模板若目錄下無任何.pr.md文件會(huì)拋出StatusCode.MEMORY_GRAPH_PROMPT_FILES_MISSING對(duì)應(yīng)的錯(cuò)誤。模板名即文件名去掉.pr.md后綴例如entity_extraction_relation_cn.pr.md注冊(cè)為entity_extraction_relation_cn。5.2 提示詞模板清單倉庫內(nèi)置了 cn/en 兩套共 22 個(gè)模板prompts/cn、prompts/en覆蓋完整抽取鏈路模板文件cn用途entity_extraction_conversation_cn.pr.md從對(duì)話消息抽取實(shí)體entity_extraction_document_cn.pr.md從文檔文本抽取實(shí)體entity_extraction_json_cn.pr.md從 JSON 內(nèi)容抽取實(shí)體entity_extraction_relation_cn.pr.md抽取事實(shí)三元組含時(shí)間信息entity_extraction_dedupe_entity_cn.pr.md實(shí)體去重entity_extraction_dedupe_relation_cn.pr.md關(guān)系去重/融合entity_extraction_entity_merge_cn.pr.md實(shí)體合并entity_extraction_relation_filter_cn.pr.md合并前的相關(guān)關(guān)系篩選entity_extraction_summary_create_cn.pr.md實(shí)體摘要與屬性生成entity_extraction_timezone_cn.pr.md時(shí)區(qū)猜測entity_extraction_check_missing_cn.pr.md缺失檢查以 entity_extraction_conversation_cn.pr.md 為例它通過#system#聲明助手角色用#user#組裝上下文模板變量包括{{source_description}}數(shù)據(jù)源描述、{{context}}歷史當(dāng)前信息、{{entity_types}}實(shí)體類型列表和{{extra_message}}可讀 schema。抽取規(guī)則強(qiáng)調(diào)對(duì)話參與者冒號(hào)前的部分必須提取為實(shí)體、代詞需消歧、動(dòng)作/關(guān)系/時(shí)間信息不得作為實(shí)體。關(guān)系抽取模板 entity_extraction_relation_cn.pr.md 則要求關(guān)系名使用英文大寫蛇形命名如FOUNDED、WORKS_AT、主體與客體必須來自實(shí)體列表且為兩個(gè)不同實(shí)體、所有時(shí)間相對(duì)于{{reference_time}}當(dāng)前 UTC 時(shí)間解析、{{tz_info}}中的時(shí)區(qū)僅供參考而非事實(shí)。六、提示詞組裝入口函數(shù)extraction_promptsopenjiuwen.core.memory.graph.extraction.extraction_prompts提供按情節(jié)類型組裝各類抽取提示詞的入口函數(shù)extraction_prompts.py所有函數(shù)統(tǒng)一返回Tuple[Dict[str, str], PromptTemplate, Dict[str, Any]]即模板變量、提示詞模板、LLM 響應(yīng)格式。6.1 實(shí)體與關(guān)系抽取extract_entity_declaration(src_type, content, history, descriptionNone, entity_typesNone, *, languagecn, extrasNone, indent2)src_typeEpisodeType情節(jié)來源類型取值為CONVERSATION/DOCUMENT/JSON對(duì)應(yīng)枚舉見 config/graph.py配置說明見 graph 配置文檔content當(dāng)前輪次或當(dāng)前文檔/JSON 內(nèi)容entity_types為None時(shí)默認(rèn)使用[EntityDef()]模板名由entity_extraction_{src_type.name.casefold()}_{language}拼接而成如entity_extraction_conversation_cn輸出模型為EntityExtraction。實(shí)現(xiàn)細(xì)節(jié)extraction_prompts.pyentity_types被格式化為0. Entity默認(rèn)實(shí)體類型。...的編號(hào)列表注入提示詞。extract_entity_attributes(entity, content, history, languagecn, extrasNone, *, indent2)為已抽取的Entity組裝摘要與屬性抽取提示模板e(cuò)ntity_extraction_summary_*輸出模型EntitySummary將entity.name、entity.content既有摘要注入模板變量若已有entity.attributes則序列化為 JSON 注入特殊處理當(dāng)實(shí)體類型為human且模板中存在summary_target時(shí)將該目標(biāo)值翻倍summary_target * 2保證人類實(shí)體獲得更充分的摘要Entity的定義參見 graph_objects。extract_relation_declaration(relation_types, entities, reference_time, tz_info, content, *, history, entity_typesNone, descriptionNone, languagecn, indent2)需傳入關(guān)系類型列表、已抽取實(shí)體聲明帶 id、參考時(shí)間戳與時(shí)區(qū)信息reference_time通過datetime.fromtimestamp(...).isoformat(timespecseconds)轉(zhuǎn)為 ISO 字符串tz_info若是 dict/list 會(huì)序列化為 JSON 字符串否則str()轉(zhuǎn)換生成id_range 1-{len(entities)}約束關(guān)系端點(diǎn)范圍輸出模型為RelationExtraction。extract_timezone(content, history, descriptionNone, languagecn, indent2)組裝時(shí)區(qū)猜測提示模板e(cuò)ntity_extraction_timezone_*輸出模型TimezonePredictions供關(guān)系抽取前異步執(zhí)行見 graph_memory/base.py 的asyncio.create_task。6.2 去重、合并與過濾函數(shù)模板輸出模型用途dedupe_entity_list(content, candidate_entities, existing_entities, entity_typesNone, history, *, descriptionNone, languagecn, indent2)dedupe_entity_*EntityDuplication判斷候選實(shí)體是否為已有實(shí)體重復(fù)項(xiàng)并給出合并 iddedupe_relation_list(content, relation, existing_relations, existing_entities, history, *, descriptionNone, languagecn, indent2)dedupe_relation_*MergeRelations判斷新關(guān)系是否與已有關(guān)系重復(fù)并給出融合結(jié)果merge_existing_entities(target, sources, languagecn, extrasNone, indent2)entity_merge_*EntitySummary將多個(gè)已有實(shí)體合并到目標(biāo)實(shí)體源實(shí)體經(jīng)format_existing_entities編號(hào)列出filter_relations_for_merge(target, relations, languagecn, extrasNone, indent2)relation_filter_*RelevantFacts針對(duì)目標(biāo)實(shí)體從候選關(guān)系列表中篩選與合并相關(guān)的 id去重細(xì)節(jié)dedupe_entity_list中已有實(shí)體從下標(biāo) 1 編號(hào)format_existing_entities(..., 1, ...)候選實(shí)體從len(existing_entities) 1開始編號(hào)format_new_entities(..., start_idx...)保證 LLM 輸出的 id 與真實(shí)庫內(nèi)下標(biāo)一一對(duì)應(yīng)。dedupe_relation_list中new_relation通過format_existing_relations([relation.model_dump()], 0).removeprefix(0. )生成去掉編號(hào)前綴后作為單條新關(guān)系展示。去重模板的判別準(zhǔn)則見 entity_extraction_dedupe_entity_cn.pr.md只有當(dāng)實(shí)體指向現(xiàn)實(shí)世界的同一事物或概念時(shí)才視為重復(fù)相關(guān)但不同或名稱相似但指向不同事物均不得標(biāo)記。關(guān)系去重模板e(cuò)ntity_extraction_dedupe_relation_cn.pr.md同樣強(qiáng)調(diào)不要融合相關(guān)但不相等的實(shí)體關(guān)系且要求combined_content為原文直接引用或簡要復(fù)述。6.3 輔助格式化函數(shù)format_new_entities(entities, entity_typesNone, start_idx1, languagecn) - str將候選實(shí)體聲明格式化為帶編號(hào)的字符串列表。若提供entity_types會(huì)先輸出類型說明{type_name}{sep}{type_description}中文分隔符為隨后以1. Alice (Person)形式列出實(shí)體中間以---分隔否則僅輸出1. Alice。對(duì)應(yīng)測試見 tests/unit_tests/core/memory/graph/extraction/test_extraction_prompts.py驗(yàn)證了空列表、起始編號(hào)偏移、類型說明輸出等行為。七、LLM 回復(fù)解析parse_responseopenjiuwen.core.memory.graph.extraction.parse_response提供從 LLM 回復(fù)中解析 JSON 與結(jié)構(gòu)化內(nèi)容的工具函數(shù)parse_response.py。parse_json(resp, output_schemaNone) - Optional[JSONLike]其中JSONLike Union[dict[str, Any], list[Any]]見 custom_types.py。解析策略分三層代碼塊優(yōu)先用正則(?s)([A-Za-z]*)\s*\n(.*?)查找 markdown 代碼塊僅解析無語言標(biāo)記或json標(biāo)注的代碼塊整段 raw_decode若無代碼塊則用JSONDecoder(strictFalse).raw_decode從響應(yīng)中[或{起始位置解碼對(duì)},結(jié)尾的截?cái)嗔斜頃?huì)先補(bǔ)}]再嘗試鍵過濾與模糊匹配若output_schema含json_schema.required則只保留這些鍵并通過difflib.get_close_matches(..., cutoff0.85)做拼寫模糊匹配try_get_key容忍 LLM 輸出與 schema 鍵名的輕微差異。同模塊還提供ensure_list(obj)將單元素 dict如{extracted_relations: [...]}解包為列表保證下游parse_all_relations拿到的是列表形態(tài)。該函數(shù)的容錯(cuò)性由單元測試覆蓋tests/unit_tests/core/memory/graph/extraction/test_parse_response.py純 JSON 對(duì)象、json代碼塊、空類型代碼塊、非 JSON 代碼塊跳過、非法 JSON 返回None、帶 required 鍵過濾、數(shù)組解析、截?cái)?JSON 等場景均有斷言。八、多語言注冊(cè)與語言校驗(yàn)抽取模塊的提示詞、schema 描述、格式化模板均以語言注冊(cè)表方式組織。cn.py與en.py分別通過register_language()將語言碼寫入REGISTERED_LANGUAGE并填充以下注冊(cè)表見 prompts/entity_extraction/base.py注冊(cè)表中文值示例英文值示例SOURCE_DESCRIPTION\n數(shù)據(jù)源描述\n{source_description}\n/數(shù)據(jù)源描述\nsource_description包裹的對(duì)應(yīng)英文模板MARK_CURRENT_MSG當(dāng)前信息\n{content}\n/當(dāng)前信息\ncurrent_messages包裹MARK_HISTORY_MSG歷史信息\n{history}\n/歷史信息\nhistory_messages包裹DISPLAY_ENTITY{i}. {name}\n{content}{i}. {name}:\n{content}RELATION_FORMAT{name}{lhs}-[{name}]-{rhs}{description}{name} ({lhs}-[{name}]-{rhs}): {description}NO_RELATION_GIVEN無Noneget_formatting_kwargs負(fù)責(zé)把 history、content 分別用當(dāng)前/歷史標(biāo)記包裹后拼入context并追加source_description與extra_message可讀 schema形成統(tǒng)一的提示詞骨架。語言參數(shù)經(jīng)ensure_valid_language(language, max_len)校驗(yàn)base.py必須是已注冊(cè)語言cn/en且長度不超過db_storage_config.language設(shè)定的上限否則拋出MEMORY_GRAPH_LANGUAGE_INVALID錯(cuò)誤。九、完整調(diào)用鏈與配置協(xié)同在GraphMemory.add_memory中抽取模塊與配置策略緊密協(xié)同graph_memory/base.pystate.prompting.language默認(rèn)cn、entity_dedupe_language、relation_extraction_language分別控制不同階段的提示詞語言對(duì)應(yīng) config/graph.py 中AddMemStrategy的chinese_entity、chinese_entity_dedupe、chinese_relation開關(guān)時(shí)區(qū)預(yù)測與關(guān)系抽取并行tz_task與關(guān)系抽取任務(wù)以asyncio.create_task并發(fā)執(zhí)行關(guān)系抽取所需的tz_info正是時(shí)區(qū)預(yù)測回復(fù)經(jīng)parse_json(..., output_schemaTimezonePredictions.response_format(...))解析后的結(jié)果關(guān)系抽取回復(fù)經(jīng)parse_all_relations見 parse_llm_response.py轉(zhuǎn)為正式Relation對(duì)象source_id/target_id為 1 基編號(hào)需減 1 映射到實(shí)體列表valid_since/valid_until由parse_iso解析為 Unix 時(shí)間戳與時(shí)區(qū)偏移同一實(shí)體自連的關(guān)系類型為EntityFact其余為RelationLLM 輸出若重復(fù)parse_all_relations會(huì)按fact內(nèi)容去重保留最長的版本。端到端的實(shí)戰(zhàn)示例見 examples/graph_memory/showcase_graph_memory.py它演示了從加載對(duì)話數(shù)據(jù)、按 chunk 調(diào)用add_memory(src_typeEpisodeType.CONVERSATION, contentchunk, user_id..., reference_time...)到實(shí)體/關(guān)系語義搜索與知識(shí)圖譜可視化的完整流程其中AddMemStrategy(summary_target100, merge_entitiesTrue, merge_relationsTrue, merge_filterTrue)直接驅(qū)動(dòng)了抽取與合并行為。文中還給出了針對(duì)不同模型的調(diào)用建議OpenAI 模型建議llm_structured_outputFalse并配合reasoning_effortminimalQwen3 系列建議llm_structured_outputTrue并配合enable_thinkingFalse。十、小結(jié)openjiuwen.core.memory.graph.extraction是 openJiuwen agent-core 圖記憶能力中承上啟下的關(guān)鍵模塊向上對(duì)接GraphMemory的寫入流水線向下通過多語言 Pydantic 模型、.pr.md提示詞模板與容錯(cuò)解析器把從非結(jié)構(gòu)化內(nèi)容到結(jié)構(gòu)化記憶這一過程封裝為可復(fù)用、可配置、可測試的 SDK 能力。理解本模塊后讀者可以基于 base.py 自定義實(shí)體/關(guān)系類型基于 extraction_prompts.py 定制抽取流程并借助 parse_response.py 建立健壯的 LLM 輸出解析鏈路。贊分享人工智能AI AgentAgent 框架大模型工具調(diào)用RAG提示工程強(qiáng)化學(xué)習(xí)【免費(fèi)下載鏈接】agent-coreopenJiuwen agent-core可提供AI Agent開發(fā)、運(yùn)行、調(diào)優(yōu)與演進(jìn)相關(guān)的全套SDK能力項(xiàng)目地址https://gitcode.com/openJiuwen/agent-core點(diǎn)擊查看免費(fèi)下載相關(guān)推薦openjiuwen agent-core 圖記憶實(shí)體與關(guān)系抽取模塊全解析多語言響應(yīng)模型、Prompt 組裝與 LLM 輸出解析openjiuwen agent core 圖記憶實(shí)體與關(guān)系抽取模塊全解析多語言響應(yīng)模型、Prompt 組裝與 LLM 輸出解析 本文以 openjiuwen人工智能AI AgentAgent 框架大模型工具調(diào)用RAG提示工程強(qiáng)化學(xué)習(xí)openJiuwen agent-core 系統(tǒng)提示詞工程實(shí)戰(zhàn)openjiuwen.harness.prompts 模塊全解openJiuwen agent core 系統(tǒng)提示詞工程實(shí)戰(zhàn)openjiuwen.harness.prompts 模塊全解 導(dǎo)讀 本指南以 prompts.人工智能AI AgentAgent 框架大模型工具調(diào)用RAG提示工程強(qiáng)化學(xué)習(xí)openjiuwen agent-core 中 Auto Harness Prompts 模塊build_auto_harness_sections 系統(tǒng)提示詞組裝機(jī)制詳解openjiuwen agent core 中 Auto Harness Prompts 模塊build_auto_harness_sections 系統(tǒng)提示人工智能AI AgentAgent 框架大模型工具調(diào)用RAG提示工程強(qiáng)化學(xué)習(xí)創(chuàng)作聲明:本文部分內(nèi)容由AI輔助生成(AIGC),僅供參考