前期准备
开始操作前,确保您已经完成了 SDK 初始化。如果您还没有初始化 SDK,请先参考快速入门文档完成。
注意:
在调用搜索接口前,需先在 腾讯云智能媒资托管控制台 开启搜索功能。
本接口不支持分页(无 marker/nextMarker),仅通过 limit 控制返回数量
本接口 QPS 使用上限为 10,不可用于业务的高频操作页面(如空间首页列表查询),如有更大 QPS 需求请提工单联系腾讯云智能媒资托管团队。
本接口(含 type=filename 基础检索)需开通白名单后使用,未开通时返回 HTTP 4xx(错误信息通常含
metainsight query is not enabled);type=filecontent 全文检索还需额外开通全文索引能力(同属白名单范畴),未开通时返回 HTTP 4xx(错误信息通常含 space does not support file content search)。搜索结果默认每页 20 个条目,可通过 limit 参数调整,取值范围 [1,100]。
是否还有更多搜索结果,不应参考 contents 的数量,而应参考 nextMarker 字段。
搜索目录与文件 - 基本检索(MI)
功能说明
searchFs 实现搜索目录与文件功能,支持 type=filename(按文件名命中)和 type=filecontent(按文件正文内容全文检索)两种子模式,并支持按关键字、文件类型、文件大小、修改时间等多种条件进行搜索,支持排序和分页。使用示例
type=filename(默认,按文件名检索)
import com.tencent.cloud.smh.ApiException;import com.tencent.cloud.smh.ApiResponse;import com.tencent.cloud.smh.api.SearchApi;import com.tencent.cloud.smh.model.SearchFsRequest;import com.tencent.cloud.smh.model.SearchFs200Response;import java.util.List;SearchFsRequest searchBody = new SearchFsRequest();searchBody.setKeywords(List.of("test", "example"));searchBody.setScope("/documents");searchBody.setInExtnames(List.of(".jpg", ".pdf"));searchBody.setFileTypes(List.of(SearchFsRequest.FileTypesEnum.FILE));searchBody.setMinFileSize(1024);searchBody.setMaxFileSize(10485760);searchBody.setType(SearchFsRequest.TypeEnum.FILENAME);try {SearchApi.APISearchFsRequest request = SearchApi.APISearchFsRequest.newBuilder().libraryId("your-library-id").spaceId("your-space-id").accessToken("your-access-token").userId("user-id").limit(20).withFavoriteStatus(1).withInode(1).searchFsRequest(searchBody).build();ApiResponse<Object> apiResponse = client.search().searchFsWithHttpInfo(request);int statusCode = apiResponse.getStatusCode();System.out.println("Status code: " + statusCode);if (statusCode == 200) {SearchFs200Response result = (SearchFs200Response) apiResponse.getData();for (var item : result.getContents()) {System.out.println("Found: " + item.getName() + " (" + item.getType() + ")");if (item.getInode() != null) {System.out.println(" Inode: " + item.getInode());}}// 如果返回了 nextMarker,说明还有更多结果,可带入下次请求继续搜索if (result.getNextMarker() != null && !result.getNextMarker().isEmpty()) {System.out.println("More results available, nextMarker: " + result.getNextMarker());}}} catch (ApiException e) {System.err.println("Error: " + e.getCode() + " - " + e.getMessage());}
type=filecontent(全文关键字检索)
注意:
此功能需联系腾讯云开通白名单后方可使用。
SearchFsRequest searchBody = new SearchFsRequest();searchBody.setKeywords(List.of("会议纪要"));searchBody.setType(SearchFsRequest.TypeEnum.FILECONTENT);try {SearchApi.APISearchFsRequest request = SearchApi.APISearchFsRequest.newBuilder().libraryId("your-library-id").spaceId("your-space-id").accessToken("your-access-token").userId("user-id").limit(10).withInode(1).searchFsRequest(searchBody).build();ApiResponse<Object> apiResponse = client.search().searchFsWithHttpInfo(request);int statusCode = apiResponse.getStatusCode();if (statusCode == 200) {SearchFs200Response result = (SearchFs200Response) apiResponse.getData();for (var item : result.getContents()) {System.out.println("Found: " + item.getName());if (item.getText() != null) {System.out.println(" Text snippet: " + item.getText() + " (page " + item.getTextPage() + ")");}if (item.getContentHighlight() != null && item.getContentHighlight().getFragments() != null) {System.out.println(" Highlight: " + item.getContentHighlight().getFragments());}}}} catch (ApiException e) {System.err.println("Error: " + e.getCode() + " - " + e.getMessage());}
参数说明
参数名 | 参数描述 | 类型 | 是否必填 |
libraryId | 媒体库 ID | String | 是 |
spaceId | 空间 ID,如果媒体库为单租户模式,则该参数固定为连字符(-);如果媒体库为多租户模式,则必须指定该参数 | String | 是 |
accessToken | 访问令牌;对于公有读媒体库或租户空间可不指定,否则需通过本参数传入或提前调用 client.withToken() 注入,二者取一 | String | 否 |
searchFsRequest | 搜索请求对象,包含详细的搜索条件 | SearchFsRequest | 否 |
userId | 用户身份识别,当访问令牌对应的权限为管理员权限且申请访问令牌时的用户身份识别为空时用来临时指定用户身份 | String | 否 |
marker | 用于顺序列出分页的标识,建议将 marker 放入请求体中传入 | String | 否 |
limit | 用于顺序列出分页时本地列出的项目数限制,取值范围 [1,100],默认值为20 | Integer | 否 |
withFavoriteStatus | 0或1,是否返回收藏状态;仅 type = filename 生效,type = filecontent 下即使传入 1 也不会返回 isFavorite | Integer | 否 |
withInode | 0或1,是否返回文件或目录 ID(inode) | Integer | 否 |
SearchFsRequest 对象说明
字段 | 类型 | 是否必填 | 说明 |
type | String | 否 | 搜索子模式,取值 filename(基础检索,按文件名命中)或 filecontent(全文关键字检索,按文件正文内容命中);默认 filename |
keywords | List<String> | 否 | 搜索关键字,字符串数组(元素间为"或"关系),数组长度上限100;type=filename 下按文件名命中,不做停用词过滤;type = filecontent 下按文件正文内容全文检索,服务端会自动过滤停用词 |
scope | String | 否 | 搜索范围,指定搜索的目录,如搜索根目录可指定为空字符串、"/"或不指定该字段;type = filecontent 下路径匹配能力有限,建议不填 |
inExtnames | List<String> | 否 | 包含的搜索文件后缀,或的关系,数组长度上限20,单元素 rune 长度上限10 |
excludeExtnames | List<String> | 否 | 不包含的搜索文件后缀,与的关系,数组长度上限20,单元素 rune 长度上限10 |
fileTypes | List<String> | 否 | 文件类型,取值 all/dir/file/symlink,或的关系 |
minFileSize | Integer | 否 | 搜索文件大小范围最小值,单位 Byte |
maxFileSize | Integer | 否 | 搜索文件大小范围最大值,单位 Byte |
modificationTimeStart | OffsetDateTime | 否 | 搜索更新时间范围起始,RFC3339 格式;若起始时间晚于结束时间返回 HTTP 4xx |
modificationTimeEnd | OffsetDateTime | 否 | 搜索更新时间范围结束,RFC3339 格式 |
orderBy | String | 否 | 排序字段;当前版本暂不支持按字段排序,字段传入不会报错但实际无效 |
orderByType | String | 否 | 排序方式,升序为 asc,降序为 desc;当前版本暂不支持,字段传入不会报错但实际无效 |
labels | List<String> | 否 | 简易文件标签,或的关系,数组长度上限 20,单元素 rune 长度上限 32 |
categories | List<String> | 否 | 文件自定义分类信息,或的关系,数组长度上限 20,取值必须属于该媒体库已声明的类目集合 |
marker | String | 否 | 分页标识 |
返回值说明
HTTP 状态码:200
搜索成功。
响应字段说明
字段 | 说明 | 类型 |
nextMarker | 用于获取后续页的分页标识,仅当 hasMore 为 true 时才返回该字段 | String |
contents | 搜索结果,可能为空数组 | List<SearchFs200ResponseContentsInner> |
Contents 数组元素说明
字段 | 说明 | 类型 | 备注 |
type | 条目类型:dir-目录或相簿;file-文件;image-图片;video-视频;symlink-符号链接 | String | |
inode | 文件或目录 ID | String | 需带 withInode=1 才会返回 |
name | 目录或相簿名或文件名 | String | |
creationTime | 创建时间或上传时间(ISO 8601) | String | |
modificationTime | 最近修改时间 | String | |
contentType | 媒体类型 | String | 仅 file 类型返回 |
versionId | 版本号 | Integer | 仅 file 类型返回 |
size | 文件大小,字符串格式避免数字精度问题 | String | 仅 file 类型返回 |
isFavorite | 是否被收藏 | Boolean | 仅 type=filename 且 withFavoriteStatus=1 时返回 |
eTag | 文件 ETag | String | 仅 file 类型返回 |
crc64 | 文件的 CRC64-ECMA182 校验值 | String | 仅 file 类型返回 |
metaData | 文件元数据信息 | Object | 仅 file 类型返回 |
userId | 创建/更新者用户 ID | String | |
previewByDoc | 是否可通过 WPS 预览 | Boolean | |
previewByCI | 是否可通过万象预览 | Boolean | |
previewAsIcon | 是否可使用预览图当做 icon | Boolean | |
fileType | 文件类别,如 doc/image/video/archive 等 | String | |
labels | 简易文件标签,字符串数组 | List<String> | |
category | 自定义文件分类,比如 image、video、doc 等 | String | |
localCreationTime | 文件对应的本地创建时间 | OffsetDateTime | |
localModificationTime | 文件对应的本地修改时间 | OffsetDateTime | |
text | 命中的正文片段 | String | 仅 type=filecontent 返回 |
textPage | 命中片段所在文档页码(整数);PDF/DOCX/PPTX 等有分页的文档才有意义,纯文本类文档为 0 | Integer | 仅 type=filecontent 返回 |
contentHighlight | 服务端侧的正文高亮片段 | Object | 仅 type=filecontent 返回 |
ContentHighlight 字段说明
字段 | 说明 | 类型 |
fragments | 高亮片段数组,关键词用 <em> 标签包裹 | List<String> |
搜索文件 - 混合检索(MI)
功能说明
searchAI 实现 AI 语义搜索功能,使用自然语言关键字做多模态语义检索。支持 type=text(文本语义搜索)和 type=pic(图片语义搜索)两种模式。此接口为高级搜索能力,使用前需联系 SMH 开发团队申请开通白名单
文本语义搜索(type=text)会对文件正文内容进行语义级别的匹配,返回与搜索语句语义相关的文档片段
图片语义搜索(type=pic)会对图片内容进行语义级别的匹配,返回与搜索语句语义相关的图片
keywords 为字符串类型(不是数组),服务端会自动做空白与特殊字符清洗,清洗后 rune 长度上限 60
若清洗后 keywords 为空,HTTP 200 返回空 contents
本接口不支持分页(无 marker/nextMarker),仅通过 limit 控制返回数量
接口 QPS 上限与具体白名单套餐挂钩,以实际开通配额为准
注意:
此功能需联系腾讯云开通白名单后方可使用。
使用示例
type=text(文本语义搜索)
import com.tencent.cloud.smh.ApiException;import com.tencent.cloud.smh.api.SearchApi;import com.tencent.cloud.smh.model.SearchAIRequest;import com.tencent.cloud.smh.model.SearchAI200Response;import java.time.OffsetDateTime;import java.util.List;SearchAIRequest searchAIBody = new SearchAIRequest();searchAIBody.setType(SearchAIRequest.TypeEnum.TEXT);searchAIBody.setKeywords("项目文档");searchAIBody.setFileTypes(List.of(SearchAIRequest.FileTypesEnum.FILE));searchAIBody.setInExtnames(List.of(".txt", ".doc", ".docx"));searchAIBody.setExcludeExtnames(List.of(".tmp"));searchAIBody.setCategories(List.of("document"));searchAIBody.setLabels(List.of("重要"));searchAIBody.setModificationTimeStart(OffsetDateTime.now().minusMonths(1));searchAIBody.setModificationTimeEnd(OffsetDateTime.now());try {SearchApi.APISearchAIRequest request = SearchApi.APISearchAIRequest.newBuilder().libraryId("your-library-id").spaceId("your-space-id").accessToken("your-access-token").userId("user-id").limit(10).searchAIRequest(searchAIBody).build();SearchAI200Response response = client.search().searchAI(request);if (response.getContents() != null) {for (var item : response.getContents()) {System.out.println("Inode: " + item.getInode() + ", Score: " + item.getScore());if (item.getText() != null) {System.out.println("Text: " + item.getText() + " (page " + item.getTextPage() + ")");}}}} catch (ApiException e) {System.err.println("Error: " + e.getCode() + " - " + e.getMessage());}
type=pic(图片语义搜索)
SearchAIRequest searchAIBody = new SearchAIRequest();searchAIBody.setType(SearchAIRequest.TypeEnum.PIC);searchAIBody.setKeywords("猫 动物");try {SearchApi.APISearchAIRequest request = SearchApi.APISearchAIRequest.newBuilder().libraryId("your-library-id").spaceId("your-space-id").accessToken("your-access-token").userId("user-id").limit(20).searchAIRequest(searchAIBody).build();SearchAI200Response response = client.search().searchAI(request);if (response.getContents() != null) {for (var item : response.getContents()) {System.out.println("Inode: " + item.getInode() + ", Score: " + item.getScore());}}} catch (ApiException e) {System.err.println("Error: " + e.getCode() + " - " + e.getMessage());}
参数说明
参数名 | 参数描述 | 类型 | 是否必填 |
libraryId | 媒体库 ID | String | 是 |
spaceId | 空间 ID,如果媒体库为单租户模式,则该参数固定为连字符(-);如果媒体库为多租户模式,则必须指定该参数 | String | 是 |
accessToken | 访问令牌;对于公有读媒体库或租户空间可不指定,否则需通过本参数传入或提前调用 client.withToken() 注入,二者取一 | String | 否 |
searchAIRequest | AI 搜索请求对象,包含详细的搜索条件 | SearchAIRequest | 是 |
userId | 用户身份识别,当访问令牌对应的权限为管理员权限且申请访问令牌时的用户身份识别为空时用来临时指定用户身份 | String | 否 |
limit | 返回最大结果数量,默认 10;type=text 取值 [0,30],type=pic 取值 [0,100] | Integer | 否 |
SearchAIRequest 对象说明
字段 | 类型 | 是否必填 | 说明 |
type | String | 是 | 子模式,取值 text(文本语义搜索)或 pic(图片语义搜索);非法值返回 HTTP 4xx |
keywords | String | 是 | 搜索语句,字符串(不是数组),服务端会自动做空白与特殊字符清洗,清洗后 rune 长度上限60;若清洗后为空则 HTTP 200 返回空 contents |
fileTypes | List<String> | 否 | 文件类型,字符串数组,取值 all/dir/file/symlink |
inExtnames | List<String> | 否 | 包含的后缀,字符串数组,数组长度上限20,单元素 rune 长度上限10 |
excludeExtnames | List<String> | 否 | 不包含的后缀,字符串数组,数组长度上限20,单元素 rune 长度上限10 |
categories | List<String> | 否 | 文件自定义分类,字符串数组,数组长度上限20,取值必须属于该媒体库已声明的类目集合 |
labels | List<String> | 否 | 简易文件标签,字符串数组,数组长度上限20,单元素 rune 长度上限32 |
modificationTimeStart | OffsetDateTime | 否 | 搜索更新时间范围起始,RFC3339 格式;若起始时间晚于结束时间返回 HTTP 4xx |
modificationTimeEnd | OffsetDateTime | 否 | 搜索更新时间范围结束,RFC3339 格式 |
返回值说明
HTTP 状态码:200
搜索成功。
响应字段说明
字段 | 说明 | 类型 |
contents | 命中结果数组,数组长度 ≤ 本次请求的 limit | List<SearchAI200ResponseContentsInner> |
Contents 数组元素说明
字段 | 说明 | 类型 | 备注 |
inode | 命中文件的 inode | String | 两种 type 均返回 |
score | 服务端计算的语义匹配得分,整数,分数越高越相关 | Integer | 两种 type 均返回 |
text | 命中文档片段(字符串) | String | 仅 type=text 返回 |
textPage | 命中片段所在页码(整数);PDF/DOCX/PPTX 等有分页的文档才有意义,纯文本类文档为 0 | Integer | 仅 type=text 返回 |
搜索聚合统计
功能说明
searchFsStats 实现搜索聚合统计功能,对搜索结果进行聚合分析,支持按文件后缀、分类、大小、媒体类型、用户 ID、文件名、文件类型等字段进行分组、计数、去重、求和、最小值、最大值、平均值等聚合操作。使用示例
import com.tencent.cloud.smh.ApiException;import com.tencent.cloud.smh.api.SearchApi;import com.tencent.cloud.smh.model.SearchFsStatsRequest;import com.tencent.cloud.smh.model.SearchFsStatsRequestAggregationsInner;import com.tencent.cloud.smh.model.SearchFsStats200Response;import java.util.List;SearchFsStatsRequestAggregationsInner aggByExt = new SearchFsStatsRequestAggregationsInner();aggByExt.setField(SearchFsStatsRequestAggregationsInner.FieldEnum.EXT_NAME);aggByExt.setOperation(SearchFsStatsRequestAggregationsInner.OperationEnum.GROUP);SearchFsStatsRequestAggregationsInner aggSizeSum = new SearchFsStatsRequestAggregationsInner();aggSizeSum.setField(SearchFsStatsRequestAggregationsInner.FieldEnum.SIZE);aggSizeSum.setOperation(SearchFsStatsRequestAggregationsInner.OperationEnum.SUM);SearchFsStatsRequest statsBody = new SearchFsStatsRequest();statsBody.setKeywords(List.of("项目"));statsBody.setScope("/");statsBody.setAggregations(List.of(aggByExt, aggSizeSum));try {SearchApi.APISearchFsStatsRequest request = SearchApi.APISearchFsStatsRequest.newBuilder().libraryId("your-library-id").spaceId("your-space-id").accessToken("your-access-token").userId("user-id").searchFsStatsRequest(statsBody).build();SearchFsStats200Response response = client.search().searchFsStats(request);if (response.getAggregations() != null) {for (var agg : response.getAggregations()) {System.out.println("Field: " + agg.getField() + ", Operation: " + agg.getOperation());if (agg.getValue() != null) {System.out.println(" Value: " + agg.getValue());}if (agg.getGroups() != null) {for (var group : agg.getGroups()) {System.out.println(" Group: " + group.getValue() + ", Count: " + group.getCount());}}}}} catch (ApiException e) {System.err.println("Error: " + e.getCode() + " - " + e.getMessage());}
参数说明
参数名 | 参数描述 | 类型 | 是否必填 |
libraryId | 媒体库 ID | String | 是 |
spaceId | 空间 ID,如果媒体库为单租户模式,则该参数固定为连字符(-);如果媒体库为多租户模式,则必须指定该参数 | String | 是 |
accessToken | 访问令牌;对于公有读媒体库或租户空间可不指定,否则需通过本参数传入或提前调用 client.withToken() 注入,二者取一 | String | 否 |
searchFsStatsRequest | 搜索聚合统计请求对象,包含搜索条件和聚合配置 | SearchFsStatsRequest | 是 |
userId | 用户身份识别,当访问令牌对应的权限为管理员权限且申请访问令牌时的用户身份识别为空时用来临时指定用户身份 | String | 否 |
SearchFsStatsRequest 对象说明
字段 | 类型 | 是否必填 | 说明 |
keywords | List<String> | 否 | 搜索关键字,字符串数组(元素间为"或"关系) |
scope | String | 否 | 搜索范围,指定搜索的目录 |
inExtnames | List<String> | 否 | 包含的搜索文件后缀,或的关系 |
excludeExtnames | List<String> | 否 | 不包含的搜索文件后缀,与的关系 |
fileTypes | List<String> | 否 | 文件类型,取值 all/dir/file/symlink |
minFileSize | Integer | 否 | 搜索文件大小范围最小值,单位 Byte |
maxFileSize | Integer | 否 | 搜索文件大小范围最大值,单位 Byte |
modificationTimeStart | OffsetDateTime | 否 | 搜索更新时间范围起始,RFC3339 格式 |
modificationTimeEnd | OffsetDateTime | 否 | 搜索更新时间范围结束,RFC3339 格式 |
labels | List<String> | 否 | 简易文件标签 |
categories | List<String> | 否 | 文件自定义分类信息 |
aggregations | List<SearchFsStatsRequestAggregationsInner> | 是 | 聚合统计数组,最多 5 个聚合项 |
SearchFsStatsRequestAggregationsInner 对象说明
字段 | 类型 | 是否必填 | 说明 |
field | String | 是 | 聚合字段:extName / category / size / contentType / userId / name / fileType |
operation | String | 是 | 聚合操作:group / count / distinct / sum / min / max / average |
subAggregations | List | 否 | 子聚合数组,仅 operation=group 时有效,最多 3 个 |
返回值说明
HTTP 状态码:200
搜索聚合统计成功。
响应字段说明
字段 | 说明 | 类型 |
isTruncated | 是否截断(即是否有更多数据未返回) | Boolean |
aggregations | 聚合结果列表 | List<SearchFsStats200ResponseAggregationsInner> |
Aggregations 数组元素说明
字段 | 说明 | 类型 |
field | 聚合字段 | String |
operation | 聚合操作 | String |
value | 聚合值(非 group 操作时返回) | BigDecimal |
groups | 分组结果数组(仅 operation=group 时返回) | List |