帮你快速理解、总结文档立即下载

搜索文件

最近更新时间:2026-09-30 16:51:31
我的收藏

前期准备

开始操作前,确保您已经完成了 SDK 初始化。如果您还没有初始化 SDK,请先参考快速入门文档完成。
注意:
在调用搜索接口前,需先在 腾讯云智能媒资托管控制台 开启搜索功能。
本接口不支持分页(无 marker/nextMarker),仅通过 limit 控制返回数量
本接口 QPS 使用上限为 10,不可用于业务的高频操作页面(如空间首页列表查询),如有更大 QPS 需求请提工单联系腾讯云智能媒资托管团队。
本接口(含 type=filename 基础检索)需开通白名单后使用,未开通时返回 HTTP 4xx(错误信息通常含 metainsight query is not enabled);type=filecontent 全文检索还需额外开通全文索引能力(同属白名单范畴),未开通时返回 HTTP 4xx(错误信息通常含 space does not support file content search)。
搜索结果默认每页 20 个条目,可通过 limit 参数调整,取值范围 [1,100]。
是否还有更多搜索结果,不应参考 contents 的数量,而应参考 nextMarker 字段。

搜索目录与文件 - 基本检索(MI)

功能说明
searchFs 实现搜索目录与文件功能,支持 type=filename(按文件名命中)和 type=filecontent(按文件正文内容全文检索)两种子模式,并支持按关键字、文件类型、文件大小、修改时间等多种条件进行搜索,支持排序和分页。
使用示例
type=filename(默认,按文件名检索)
import com.tencent.cloud.smh.ApiException;
import com.tencent.cloud.smh.ApiResponse;
import com.tencent.cloud.smh.api.SearchApi;
import com.tencent.cloud.smh.model.SearchFsRequest;
import com.tencent.cloud.smh.model.SearchFs200Response;
import java.util.List;

SearchFsRequest searchBody = new SearchFsRequest();
searchBody.setKeywords(List.of("test", "example"));
searchBody.setScope("/documents");
searchBody.setInExtnames(List.of(".jpg", ".pdf"));
searchBody.setFileTypes(List.of(SearchFsRequest.FileTypesEnum.FILE));
searchBody.setMinFileSize(1024);
searchBody.setMaxFileSize(10485760);
searchBody.setType(SearchFsRequest.TypeEnum.FILENAME);

try {
SearchApi.APISearchFsRequest request = SearchApi.APISearchFsRequest.newBuilder()
.libraryId("your-library-id")
.spaceId("your-space-id")
.accessToken("your-access-token")
.userId("user-id")
.limit(20)
.withFavoriteStatus(1)
.withInode(1)
.searchFsRequest(searchBody)
.build();

ApiResponse<Object> apiResponse = client.search().searchFsWithHttpInfo(request);
int statusCode = apiResponse.getStatusCode();
System.out.println("Status code: " + statusCode);

if (statusCode == 200) {
SearchFs200Response result = (SearchFs200Response) apiResponse.getData();
for (var item : result.getContents()) {
System.out.println("Found: " + item.getName() + " (" + item.getType() + ")");
if (item.getInode() != null) {
System.out.println(" Inode: " + item.getInode());
}
}

// 如果返回了 nextMarker,说明还有更多结果,可带入下次请求继续搜索
if (result.getNextMarker() != null && !result.getNextMarker().isEmpty()) {
System.out.println("More results available, nextMarker: " + result.getNextMarker());
}
}
} catch (ApiException e) {
System.err.println("Error: " + e.getCode() + " - " + e.getMessage());
}
type=filecontent(全文关键字检索)
注意:
此功能需联系腾讯云开通白名单后方可使用。
SearchFsRequest searchBody = new SearchFsRequest();
searchBody.setKeywords(List.of("会议纪要"));
searchBody.setType(SearchFsRequest.TypeEnum.FILECONTENT);

try {
SearchApi.APISearchFsRequest request = SearchApi.APISearchFsRequest.newBuilder()
.libraryId("your-library-id")
.spaceId("your-space-id")
.accessToken("your-access-token")
.userId("user-id")
.limit(10)
.withInode(1)
.searchFsRequest(searchBody)
.build();

ApiResponse<Object> apiResponse = client.search().searchFsWithHttpInfo(request);
int statusCode = apiResponse.getStatusCode();

if (statusCode == 200) {
SearchFs200Response result = (SearchFs200Response) apiResponse.getData();
for (var item : result.getContents()) {
System.out.println("Found: " + item.getName());
if (item.getText() != null) {
System.out.println(" Text snippet: " + item.getText() + " (page " + item.getTextPage() + ")");
}
if (item.getContentHighlight() != null && item.getContentHighlight().getFragments() != null) {
System.out.println(" Highlight: " + item.getContentHighlight().getFragments());
}
}
}
} catch (ApiException e) {
System.err.println("Error: " + e.getCode() + " - " + e.getMessage());
}
参数说明
参数名
参数描述
类型
是否必填
libraryId
媒体库 ID
String
是
spaceId
空间 ID,如果媒体库为单租户模式,则该参数固定为连字符(-);如果媒体库为多租户模式,则必须指定该参数
String
是
accessToken
访问令牌;对于公有读媒体库或租户空间可不指定,否则需通过本参数传入或提前调用 client.withToken() 注入,二者取一
String
否
searchFsRequest
搜索请求对象,包含详细的搜索条件
SearchFsRequest
否
userId
用户身份识别,当访问令牌对应的权限为管理员权限且申请访问令牌时的用户身份识别为空时用来临时指定用户身份
String
否
marker
用于顺序列出分页的标识,建议将 marker 放入请求体中传入
String
否
limit
用于顺序列出分页时本地列出的项目数限制,取值范围 [1,100],默认值为20
Integer
否
withFavoriteStatus
0或1,是否返回收藏状态;仅 type = filename 生效,type = filecontent 下即使传入 1 也不会返回 isFavorite
Integer
否
withInode
0或1,是否返回文件或目录 ID(inode)
Integer
否
SearchFsRequest 对象说明
字段
类型
是否必填
说明
type
String
否
搜索子模式,取值 filename(基础检索,按文件名命中)或 filecontent(全文关键字检索,按文件正文内容命中);默认 filename
keywords
List<String>
否
搜索关键字,字符串数组(元素间为"或"关系),数组长度上限100;type=filename 下按文件名命中,不做停用词过滤;type = filecontent 下按文件正文内容全文检索,服务端会自动过滤停用词
scope
String
否
搜索范围,指定搜索的目录,如搜索根目录可指定为空字符串、"/"或不指定该字段;type = filecontent 下路径匹配能力有限,建议不填
inExtnames
List<String>
否
包含的搜索文件后缀,或的关系,数组长度上限20,单元素 rune 长度上限10
excludeExtnames
List<String>
否
不包含的搜索文件后缀,与的关系,数组长度上限20,单元素 rune 长度上限10
fileTypes
List<String>
否
文件类型,取值 all/dir/file/symlink,或的关系
minFileSize
Integer
否
搜索文件大小范围最小值,单位 Byte
maxFileSize
Integer
否
搜索文件大小范围最大值,单位 Byte
modificationTimeStart
OffsetDateTime
否
搜索更新时间范围起始,RFC3339 格式;若起始时间晚于结束时间返回 HTTP 4xx
modificationTimeEnd
OffsetDateTime
否
搜索更新时间范围结束,RFC3339 格式
orderBy
String
否
排序字段;当前版本暂不支持按字段排序,字段传入不会报错但实际无效
orderByType
String
否
排序方式,升序为 asc,降序为 desc;当前版本暂不支持,字段传入不会报错但实际无效
labels
List<String>
否
简易文件标签,或的关系,数组长度上限 20,单元素 rune 长度上限 32
categories
List<String>
否
文件自定义分类信息,或的关系,数组长度上限 20,取值必须属于该媒体库已声明的类目集合
marker
String
否
分页标识
返回值说明
HTTP 状态码:200
搜索成功。
响应字段说明
字段
说明
类型
nextMarker
用于获取后续页的分页标识,仅当 hasMore 为 true 时才返回该字段
String
contents
搜索结果,可能为空数组
List<SearchFs200ResponseContentsInner>
Contents 数组元素说明
字段
说明
类型
备注
type
条目类型:dir-目录或相簿;file-文件;image-图片;video-视频;symlink-符号链接
String

inode
文件或目录 ID
String
需带 withInode=1 才会返回
name
目录或相簿名或文件名
String

creationTime
创建时间或上传时间(ISO 8601)
String

modificationTime
最近修改时间
String

contentType
媒体类型
String
仅 file 类型返回
versionId
版本号
Integer
仅 file 类型返回
size
文件大小,字符串格式避免数字精度问题
String
仅 file 类型返回
isFavorite
是否被收藏
Boolean
仅 type=filename 且 withFavoriteStatus=1 时返回
eTag
文件 ETag
String
仅 file 类型返回
crc64
文件的 CRC64-ECMA182 校验值
String
仅 file 类型返回
metaData
文件元数据信息
Object
仅 file 类型返回
userId
创建/更新者用户 ID
String

previewByDoc
是否可通过 WPS 预览
Boolean

previewByCI
是否可通过万象预览
Boolean

previewAsIcon
是否可使用预览图当做 icon
Boolean

fileType
文件类别,如 doc/image/video/archive 等
String

labels
简易文件标签,字符串数组
List<String>

category
自定义文件分类,比如 image、video、doc 等
String

localCreationTime
文件对应的本地创建时间
OffsetDateTime

localModificationTime
文件对应的本地修改时间
OffsetDateTime

text
命中的正文片段
String
仅 type=filecontent 返回
textPage
命中片段所在文档页码(整数);PDF/DOCX/PPTX 等有分页的文档才有意义,纯文本类文档为 0
Integer
仅 type=filecontent 返回
contentHighlight
服务端侧的正文高亮片段
Object
仅 type=filecontent 返回
ContentHighlight 字段说明
字段
说明
类型
fragments
高亮片段数组,关键词用 <em> 标签包裹
List<String>

搜索文件 - 混合检索(MI)

功能说明
searchAI 实现 AI 语义搜索功能,使用自然语言关键字做多模态语义检索。支持 type=text(文本语义搜索)和 type=pic(图片语义搜索)两种模式。
此接口为高级搜索能力,使用前需联系 SMH 开发团队申请开通白名单
文本语义搜索(type=text)会对文件正文内容进行语义级别的匹配,返回与搜索语句语义相关的文档片段
图片语义搜索(type=pic)会对图片内容进行语义级别的匹配,返回与搜索语句语义相关的图片
keywords 为字符串类型(不是数组),服务端会自动做空白与特殊字符清洗,清洗后 rune 长度上限 60
若清洗后 keywords 为空,HTTP 200 返回空 contents
本接口不支持分页(无 marker/nextMarker),仅通过 limit 控制返回数量
接口 QPS 上限与具体白名单套餐挂钩,以实际开通配额为准
注意:
此功能需联系腾讯云开通白名单后方可使用。
使用示例
type=text(文本语义搜索)
import com.tencent.cloud.smh.ApiException;
import com.tencent.cloud.smh.api.SearchApi;
import com.tencent.cloud.smh.model.SearchAIRequest;
import com.tencent.cloud.smh.model.SearchAI200Response;

import java.time.OffsetDateTime;
import java.util.List;

SearchAIRequest searchAIBody = new SearchAIRequest();
searchAIBody.setType(SearchAIRequest.TypeEnum.TEXT);
searchAIBody.setKeywords("项目文档");
searchAIBody.setFileTypes(List.of(SearchAIRequest.FileTypesEnum.FILE));
searchAIBody.setInExtnames(List.of(".txt", ".doc", ".docx"));
searchAIBody.setExcludeExtnames(List.of(".tmp"));
searchAIBody.setCategories(List.of("document"));
searchAIBody.setLabels(List.of("重要"));
searchAIBody.setModificationTimeStart(OffsetDateTime.now().minusMonths(1));
searchAIBody.setModificationTimeEnd(OffsetDateTime.now());

try {
SearchApi.APISearchAIRequest request = SearchApi.APISearchAIRequest.newBuilder()
.libraryId("your-library-id")
.spaceId("your-space-id")
.accessToken("your-access-token")
.userId("user-id")
.limit(10)
.searchAIRequest(searchAIBody)
.build();

SearchAI200Response response = client.search().searchAI(request);
if (response.getContents() != null) {
for (var item : response.getContents()) {
System.out.println("Inode: " + item.getInode() + ", Score: " + item.getScore());
if (item.getText() != null) {
System.out.println("Text: " + item.getText() + " (page " + item.getTextPage() + ")");
}
}
}
} catch (ApiException e) {
System.err.println("Error: " + e.getCode() + " - " + e.getMessage());
}
type=pic(图片语义搜索)
SearchAIRequest searchAIBody = new SearchAIRequest();
searchAIBody.setType(SearchAIRequest.TypeEnum.PIC);
searchAIBody.setKeywords("猫 动物");

try {
SearchApi.APISearchAIRequest request = SearchApi.APISearchAIRequest.newBuilder()
.libraryId("your-library-id")
.spaceId("your-space-id")
.accessToken("your-access-token")
.userId("user-id")
.limit(20)
.searchAIRequest(searchAIBody)
.build();

SearchAI200Response response = client.search().searchAI(request);
if (response.getContents() != null) {
for (var item : response.getContents()) {
System.out.println("Inode: " + item.getInode() + ", Score: " + item.getScore());
}
}
} catch (ApiException e) {
System.err.println("Error: " + e.getCode() + " - " + e.getMessage());
}
参数说明
参数名
参数描述
类型
是否必填
libraryId
媒体库 ID
String
是
spaceId
空间 ID,如果媒体库为单租户模式,则该参数固定为连字符(-);如果媒体库为多租户模式,则必须指定该参数
String
是
accessToken
访问令牌;对于公有读媒体库或租户空间可不指定,否则需通过本参数传入或提前调用 client.withToken() 注入,二者取一
String
否
searchAIRequest
AI 搜索请求对象,包含详细的搜索条件
SearchAIRequest
是
userId
用户身份识别,当访问令牌对应的权限为管理员权限且申请访问令牌时的用户身份识别为空时用来临时指定用户身份
String
否
limit
返回最大结果数量,默认 10;type=text 取值 [0,30],type=pic 取值 [0,100]
Integer
否
SearchAIRequest 对象说明
字段
类型
是否必填
说明
type
String
是
子模式,取值 text(文本语义搜索)或 pic(图片语义搜索);非法值返回 HTTP 4xx
keywords
String
是
搜索语句,字符串(不是数组),服务端会自动做空白与特殊字符清洗,清洗后 rune 长度上限60;若清洗后为空则 HTTP 200 返回空 contents
fileTypes
List<String>
否
文件类型,字符串数组,取值 all/dir/file/symlink
inExtnames
List<String>
否
包含的后缀,字符串数组,数组长度上限20,单元素 rune 长度上限10
excludeExtnames
List<String>
否
不包含的后缀,字符串数组,数组长度上限20,单元素 rune 长度上限10
categories
List<String>
否
文件自定义分类,字符串数组,数组长度上限20,取值必须属于该媒体库已声明的类目集合
labels
List<String>
否
简易文件标签,字符串数组,数组长度上限20,单元素 rune 长度上限32
modificationTimeStart
OffsetDateTime
否
搜索更新时间范围起始,RFC3339 格式;若起始时间晚于结束时间返回 HTTP 4xx
modificationTimeEnd
OffsetDateTime
否
搜索更新时间范围结束,RFC3339 格式
返回值说明
HTTP 状态码:200
搜索成功。
响应字段说明
字段
说明
类型
contents
命中结果数组,数组长度 ≤ 本次请求的 limit
List<SearchAI200ResponseContentsInner>
Contents 数组元素说明
字段
说明
类型
备注
inode
命中文件的 inode
String
两种 type 均返回
score
服务端计算的语义匹配得分,整数,分数越高越相关
Integer
两种 type 均返回
text
命中文档片段(字符串)
String
仅 type=text 返回
textPage
命中片段所在页码(整数);PDF/DOCX/PPTX 等有分页的文档才有意义,纯文本类文档为 0
Integer
仅 type=text 返回

搜索聚合统计

功能说明
searchFsStats 实现搜索聚合统计功能,对搜索结果进行聚合分析,支持按文件后缀、分类、大小、媒体类型、用户 ID、文件名、文件类型等字段进行分组、计数、去重、求和、最小值、最大值、平均值等聚合操作。
使用示例
import com.tencent.cloud.smh.ApiException;
import com.tencent.cloud.smh.api.SearchApi;
import com.tencent.cloud.smh.model.SearchFsStatsRequest;
import com.tencent.cloud.smh.model.SearchFsStatsRequestAggregationsInner;
import com.tencent.cloud.smh.model.SearchFsStats200Response;
import java.util.List;

SearchFsStatsRequestAggregationsInner aggByExt = new SearchFsStatsRequestAggregationsInner();
aggByExt.setField(SearchFsStatsRequestAggregationsInner.FieldEnum.EXT_NAME);
aggByExt.setOperation(SearchFsStatsRequestAggregationsInner.OperationEnum.GROUP);

SearchFsStatsRequestAggregationsInner aggSizeSum = new SearchFsStatsRequestAggregationsInner();
aggSizeSum.setField(SearchFsStatsRequestAggregationsInner.FieldEnum.SIZE);
aggSizeSum.setOperation(SearchFsStatsRequestAggregationsInner.OperationEnum.SUM);

SearchFsStatsRequest statsBody = new SearchFsStatsRequest();
statsBody.setKeywords(List.of("项目"));
statsBody.setScope("/");
statsBody.setAggregations(List.of(aggByExt, aggSizeSum));

try {
SearchApi.APISearchFsStatsRequest request = SearchApi.APISearchFsStatsRequest.newBuilder()
.libraryId("your-library-id")
.spaceId("your-space-id")
.accessToken("your-access-token")
.userId("user-id")
.searchFsStatsRequest(statsBody)
.build();

SearchFsStats200Response response = client.search().searchFsStats(request);
if (response.getAggregations() != null) {
for (var agg : response.getAggregations()) {
System.out.println("Field: " + agg.getField() + ", Operation: " + agg.getOperation());
if (agg.getValue() != null) {
System.out.println(" Value: " + agg.getValue());
}
if (agg.getGroups() != null) {
for (var group : agg.getGroups()) {
System.out.println(" Group: " + group.getValue() + ", Count: " + group.getCount());
}
}
}
}
} catch (ApiException e) {
System.err.println("Error: " + e.getCode() + " - " + e.getMessage());
}
参数说明
参数名
参数描述
类型
是否必填
libraryId
媒体库 ID
String
是
spaceId
空间 ID,如果媒体库为单租户模式,则该参数固定为连字符(-);如果媒体库为多租户模式,则必须指定该参数
String
是
accessToken
访问令牌;对于公有读媒体库或租户空间可不指定,否则需通过本参数传入或提前调用 client.withToken() 注入,二者取一
String
否
searchFsStatsRequest
搜索聚合统计请求对象,包含搜索条件和聚合配置
SearchFsStatsRequest
是
userId
用户身份识别,当访问令牌对应的权限为管理员权限且申请访问令牌时的用户身份识别为空时用来临时指定用户身份
String
否
SearchFsStatsRequest 对象说明
字段
类型
是否必填
说明
keywords
List<String>
否
搜索关键字,字符串数组(元素间为"或"关系)
scope
String
否
搜索范围,指定搜索的目录
inExtnames
List<String>
否
包含的搜索文件后缀,或的关系
excludeExtnames
List<String>
否
不包含的搜索文件后缀,与的关系
fileTypes
List<String>
否
文件类型,取值 all/dir/file/symlink
minFileSize
Integer
否
搜索文件大小范围最小值,单位 Byte
maxFileSize
Integer
否
搜索文件大小范围最大值,单位 Byte
modificationTimeStart
OffsetDateTime
否
搜索更新时间范围起始,RFC3339 格式
modificationTimeEnd
OffsetDateTime
否
搜索更新时间范围结束,RFC3339 格式
labels
List<String>
否
简易文件标签
categories
List<String>
否
文件自定义分类信息
aggregations
List<SearchFsStatsRequestAggregationsInner>
是
聚合统计数组,最多 5 个聚合项
SearchFsStatsRequestAggregationsInner 对象说明
字段
类型
是否必填
说明
field
String
是
聚合字段:extName / category / size / contentType / userId / name / fileType
operation
String
是
聚合操作:group / count / distinct / sum / min / max / average
subAggregations
List
否
子聚合数组,仅 operation=group 时有效,最多 3 个
返回值说明
HTTP 状态码:200
搜索聚合统计成功。
响应字段说明
字段
说明
类型
isTruncated
是否截断(即是否有更多数据未返回)
Boolean
aggregations
聚合结果列表
List<SearchFsStats200ResponseAggregationsInner>
Aggregations 数组元素说明
字段
说明
类型
field
聚合字段
String
operation
聚合操作
String
value
聚合值(非 group 操作时返回)
BigDecimal
groups
分组结果数组(仅 operation=group 时返回)
List