↑↓ 选择↵ 打开⌫ 切换范围完整搜索

PG.CENTER 连接 PostgreSQL 文档、百科与生态知识。由 Pigsty 维护。

支持中的版本: 当前版本 (18) / 17 / 16 / 15 / 14
开发中的版本: 19 / 20devel
已结束支持的版本: 13 / 12 / 11 / 10 / 9.6 / 9.5 / 9.4 / 9.3 / 9.2 / 9.1 / 9.0 / 8.4 / 8.3

9.13. 文本检索函数和操作符 #

表 9.41、表 9.42 以及表 9.43 总结了为全文检索提供的函数和操作符。PostgreSQL 的文本检索功能的详细解释可参考第 12 章。

表 9.41. 文本检索操作符

操作符

描述

示例

tsvector @@ tsquery → boolean

tsquery @@ tsvector → boolean

tsvector 匹配 tsquery 吗?(参数可以按任意顺序给出。)

to_tsvector('fat cats ate rats') @@ to_tsquery('cat & rat') → t

text @@ tsquery → boolean

隐式调用 to_tsvector() 后,文本字符串是否匹配 tsquery?

'fat cats ate rats' @@ to_tsquery('cat & rat') → t

tsvector @@@ tsquery → boolean

tsquery @@@ tsvector → boolean

这是@@已弃用的同义词。

to_tsvector('fat cats ate rats') @@@ to_tsquery('cat & rat') → t

tsvector || tsvector → tsvector

连接两个 tsvector。如果两个输入都包含词位位置,则相应地调整第二个输入的位置。

'a:1 b:2'::tsvector || 'c:1 d:2 b:3'::tsvector → 'a':1 'b':2,5 'c':3 'd':4

tsquery && tsquery → tsquery

将两个 tsquery 做 AND 组合,生成一个匹配同时满足两个输入查询的文档的查询。

'fat | rat'::tsquery && 'cat'::tsquery → ( 'fat' | 'rat' ) & 'cat'

tsquery || tsquery → tsquery

将两个 tsquery 做 OR 组合,生成一个匹配任一输入查询的文档的查询。

'fat | rat'::tsquery || 'cat'::tsquery → 'fat' | 'rat' | 'cat'

!! tsquery → tsquery

对 tsquery 取反,生成匹配那些不满足输入查询的文档的查询。

!! 'cat'::tsquery → !'cat'

tsquery <-> tsquery → tsquery

构造一个短语查询;当两个输入查询分别匹配相邻的词位时,该查询匹配。

to_tsquery('fat') <-> to_tsquery('rat') → 'fat' <-> 'rat'

tsquery @> tsquery → boolean

第一个 tsquery 包含了第二个吗?(这只考虑出现在一个查询中的所有词位是否出现在另一个查询中,忽略了组合操作符。)

'cat'::tsquery @> 'cat & rat'::tsquery → f

tsquery <@ tsquery → boolean

第一个 tsquery 包含在第二个中吗?(这只考虑出现在一个查询中的所有词位是否出现在另一个查询中,而忽略了组合操作符。)

'cat'::tsquery <@ 'cat & rat'::tsquery → t

'cat'::tsquery <@ '!cat & rat'::tsquery → t


除了这些专用操作符之外,表 9.1 中所示的常用比较操作符也适用于 tsvector 和 tsquery 类型。这些操作符对文本检索用处不大,但可以用于其他用途,例如在这些类型的列上建立唯一索引。

表 9.42. 文本检索函数

函数

描述

示例

array_to_tsvector ( text[] ) → tsvector

将词位数组转换为 tsvector。给定的字符串按原样使用,不做进一步处理。

array_to_tsvector('{fat,cat,rat}'::text[]) → 'cat' 'fat' 'rat'

get_current_ts_config ( ) → regconfig

返回当前默认文本检索配置的 OID(由 default_text_search_config 设置)。

get_current_ts_config() → english

length ( tsvector ) → integer

返回 tsvector 中的词位数。

length('fat:2,4 cat:3 rat:5A'::tsvector) → 3

numnode ( tsquery ) → integer

返回 tsquery 中词位和操作符的数目。

numnode('(fat & rat) | cat'::tsquery) → 5

plainto_tsquery ( [ config regconfig, ] query text ) → tsquery

将文本转换为 tsquery,根据指定的或默认配置对单词进行正规化。字符串中的任何标点符号都会被忽略(它们不决定查询操作符)。生成的查询匹配包含文本中所有非停用词的文档。

plainto_tsquery('english', 'The Fat Rats') → 'fat' & 'rat'

phraseto_tsquery ( [ config regconfig, ] query text ) → tsquery

将文本转换为 tsquery,根据指定的或默认配置对单词进行正规化。字符串中的任何标点符号都会被忽略(它不决定查询操作符)。结果查询匹配包含文本中所有非停用词的短语。

phraseto_tsquery('english', 'The Fat Rats') → 'fat' <-> 'rat'

phraseto_tsquery('english', 'The Cat and Rats') → 'cat' <2> 'rat'

websearch_to_tsquery ( [ config regconfig, ] query text ) → tsquery

将文本转换为 tsquery,根据指定的或默认配置对单词进行正规化。带引号的单词序列被转换为短语测试。“or”一词产生 OR 操作符,短横线产生 NOT 操作符;其他标点符号会被忽略。这类似于一些常见网络搜索工具的行为。

websearch_to_tsquery('english', '"fat rat" or cat dog') → 'fat' <-> 'rat' | 'cat' & 'dog'

querytree ( tsquery ) → text

生成 tsquery 中可索引部分的表示。结果为空或仅为 T 表示该查询不可索引。

querytree('foo & ! bar'::tsquery) → 'foo'

setweight ( vector tsvector, weight "char" ) → tsvector

将指定的 weight 赋给 vector 的每个元素。

setweight('fat:2,4 cat:3 rat:5B'::tsvector, 'A') → 'cat':3A 'fat':2A,4A 'rat':5A

setweight ( vector tsvector, weight "char", lexemes text[] ) → tsvector

为 vector 中列在 lexemes 内的元素赋予指定的 weight。

setweight('fat:2,4 cat:3 rat:5,6B'::tsvector, 'A', '{cat,rat}') → 'cat':3A 'fat':2,4 'rat':5A,6A

strip ( tsvector ) → tsvector

从 tsvector 中移除位置和权重。

strip('fat:2,4 cat:3 rat:5A'::tsvector) → 'cat' 'fat' 'rat'

to_tsquery ( [ config regconfig, ] query text ) → tsquery

将文本转换为 tsquery,根据指定的或默认配置对单词进行正规化。单词必须由有效的 tsquery 操作符组合。

to_tsquery('english', 'The & Fat & Rats') → 'fat' & 'rat'

to_tsvector ( [ config regconfig, ] document text ) → tsvector

将文本转换为 tsvector,根据指定的或默认配置对单词进行正规化。结果中包含位置信息。

to_tsvector('english', 'The Fat Rats') → 'fat':2 'rat':3

to_tsvector ( [ config regconfig, ] document json ) → tsvector

to_tsvector ( [ config regconfig, ] document jsonb ) → tsvector

将 JSON 文档中的每个字符串值转换为 tsvector,根据指定的或默认配置对单词进行正规化。然后将结果按文档顺序连接起来以产生输出。生成位置信息时,视为每对字符串值之间存在一个停用词。(注意,当输入为 jsonb 时,JSON 对象字段的“文档顺序”取决于具体实现;请注意这些示例中的差异。)

to_tsvector('english', '{"aa": "The Fat Rats", "b": "dog"}'::json) → 'dog':5 'fat':2 'rat':3

to_tsvector('english', '{"aa": "The Fat Rats", "b": "dog"}'::jsonb) → 'dog':1 'fat':4 'rat':5

json_to_tsvector ( [ config regconfig, ] document json, filter jsonb ) → tsvector

jsonb_to_tsvector ( [ config regconfig, ] document jsonb, filter jsonb ) → tsvector

选择 filter 请求的 JSON 文档中的每个项,并将每个项转换为 tsvector,根据指定的或默认配置对单词进行正规化。然后将结果按文档顺序连接起来以产生输出。位置信息就像在每对选定的项目之间存在一个停用词一样生成。(注意,当输入为 jsonb 时,JSON 对象字段的“文档顺序”取决于实现。)filter 必须是一个 jsonb 数组,其中包含 0 个或多个关键字:"string"(包括所有字符串值),"numeric"(包括所有数值),"boolean"(包括所有布尔值),"key"(包括所有键),或 "all"(包括以上所有内容)。作为一种特殊情况,该 filter 也可以是这些关键字之一的简单 JSON 值。

json_to_tsvector('english', '{"a": "The Fat Rats", "b": 123}'::json, '["string", "numeric"]') → '123':5 'fat':2 'rat':3

json_to_tsvector('english', '{"cat": "The Fat Rats", "dog": 123}'::json, '"all"') → '123':9 'cat':1 'dog':7 'fat':4 'rat':5

ts_delete ( vector tsvector, lexeme text ) → tsvector

从 vector 中删除给定的 lexeme 的所有出现。

ts_delete('fat:2,4 cat:3 rat:5A'::tsvector, 'fat') → 'cat':3 'rat':5A

ts_delete ( vector tsvector, lexemes text[] ) → tsvector

从 vector 中删除 lexemes 所列词位的所有出现。

ts_delete('fat:2,4 cat:3 rat:5A'::tsvector, ARRAY['fat','rat']) → 'cat':3

ts_filter ( vector tsvector, weights "char"[] ) → tsvector

只从 vector 中选择具有给定 weights 的元素。

ts_filter('fat:2,4 cat:3b,7c rat:5A'::tsvector, '{a,b}') → 'cat':3B 'rat':5A

ts_headline ( [ config regconfig, ] document text, query tsquery [, options text ] ) → text

以缩略形式显示 query 在 document 中的匹配项;后者必须是原始文本,不能是 tsvector。在匹配查询之前,文档中的单词将根据指定的或默认配置进行正规化。第 12.3.4 节中讨论了该函数的使用,还描述了可用的 options。

ts_headline('The fat cat ate the rat.', 'cat') → The fat <b>cat</b> ate the rat.

ts_headline ( [ config regconfig, ] document json, query tsquery [, options text ] ) → text

ts_headline ( [ config regconfig, ] document jsonb, query tsquery [, options text ] ) → text

以缩略形式显示 query 在 JSON document 字符串值中的匹配项。更多细节请参见第 12.3.4 节。

ts_headline('{"cat":"raining cats and dogs"}'::jsonb, 'cat') → {"cat": "raining <b>cats</b> and dogs"}

ts_rank ( [ weights real[], ] vector tsvector, query tsquery [, normalization integer ] ) → real

计算一个分数,显示 vector 与 query 的匹配程度。详情请参见第 12.3.3 节。

ts_rank(to_tsvector('raining cats and dogs'), 'cat') → 0.06079271

ts_rank_cd ( [ weights real[], ] vector tsvector, query tsquery [, normalization integer ] ) → real

使用覆盖密度算法计算一个分数,显示 vector 与 query 的匹配程度。详情参见第 12.3.3 节。

ts_rank_cd(to_tsvector('raining cats and dogs'), 'cat') → 0.1

ts_rewrite ( query tsquery, target tsquery, substitute tsquery ) → tsquery

在 query 中使用 substitute 替换出现的 target。详情参见第 12.4.2.1 节。

ts_rewrite('a & b'::tsquery, 'a'::tsquery, 'foo|bar'::tsquery) → 'b' & ( 'foo' | 'bar' )

ts_rewrite ( query tsquery, select text ) → tsquery

根据执行 SELECT 命令得到的目标和替换项,替换 query 中的相应部分。详情参见第 12.4.2.1 节。

SELECT ts_rewrite('a & b'::tsquery, 'SELECT t,s FROM aliases') → 'b' & ( 'foo' | 'bar' )

tsquery_phrase ( query1 tsquery, query2 tsquery ) → tsquery

构造一个短语查询,在连续的词位上搜索 query1 和 query2 的匹配项(与<-> 操作符相同)。

tsquery_phrase(to_tsquery('fat'), to_tsquery('cat')) → 'fat' <-> 'cat'

tsquery_phrase ( query1 tsquery, query2 tsquery, distance integer ) → tsquery

构造一个短语查询,用于搜索 query1 和 query2 的匹配项,其匹配位置恰好相距 distance 个词位。

tsquery_phrase(to_tsquery('fat'), to_tsquery('cat'), 10) → 'fat' <10> 'cat'

tsvector_to_array ( tsvector ) → text[]

将 tsvector 转换为词位的数组。

tsvector_to_array('fat:2,4 cat:3 rat:5A'::tsvector) → {cat,fat,rat}

unnest ( tsvector ) → setof record ( lexeme text, positions smallint[], weights text )

将 tsvector 展开为一组行,每行对应一个词位。

select * from unnest('cat:3 fat:2,4 rat:5A'::tsvector) →

 lexeme | positions | weights
--------+-----------+---------
 cat    | {3}       | {D}
 fat    | {2,4}     | {D,D}
 rat    | {5}       | {A}


注意

所有接受一个可选的 regconfig 参数的文本检索函数在省略该参数时,会使用由 default_text_search_config 指定的配置。

表 9.43 中的函数被单独列出,因为它们通常不被用于日常的文本检索操作。它们主要有助于开发和调试新的文本检索配置。

表 9.43. 文本检索调试函数

函数

描述

示例

ts_debug ( [ config regconfig, ] document text ) → setof record ( alias text, description text, token text, dictionaries regdictionary[], dictionary regdictionary, lexemes text[] )

根据指定的或默认的文本检索配置从 document 中提取和正规化词元,并返回关于每个词元是如何处理的信息。详情参见第 12.8.1 节。

ts_debug('english', 'The Brightest supernovaes') → (asciiword,"Word, all ASCII",The,{english_stem},english_stem,{}) ...

ts_lexize ( dict regdictionary, token text ) → text[]

如果词典识别输入词元,则返回由替换词位组成的数组;如果词典识别该词元,但它是停用词,则返回空数组;如果词典无法识别该词元,则返回 NULL。详情参见第 12.8.3 节。

ts_lexize('english_stem', 'stars') → {star}

ts_parse ( parser_name text, document text ) → setof record ( tokid integer, token text )

使用指定名称的解析器从 document 中提取词元。详情参见第 12.8.2 节。

ts_parse('default', 'foo - bar') → (1,foo) ...

ts_parse ( parser_oid oid, document text ) → setof record ( tokid integer, token text )

使用 OID 指定的解析器从 document 中提取词元。详情参见第 12.8.2 节。

ts_parse(3722, 'foo - bar') → (1,foo) ...

ts_token_type ( parser_name text ) → setof record ( tokid integer, alias text, description text )

返回一个表,该表描述指定名称的解析器可以识别的每种类型的词元。详情参见第 12.8.2 节。

ts_token_type('default') → (1,asciiword,"Word, all ASCII") ...

ts_token_type ( parser_oid oid ) → setof record ( tokid integer, alias text, description text )

返回一个表,该表描述 OID 指定的解析器可以识别的每种词元类型。详情参见第 12.8.2 节。

ts_token_type(3722) → (1,asciiword,"Word, all ASCII") ...

ts_stat ( sqlquery text [, weights text ] ) → setof record ( word text, ndoc integer, nentry integer )

执行 sqlquery,该查询必须返回单个 tsvector 列,并返回数据中每个不同词位的统计信息。详情参见第 12.4.4 节。

ts_stat('SELECT vector FROM apod') → (foo,10,15) ...


报告文档问题

阅读 上游文档. 通过 PostgreSQL 文档反馈表单.