原厂 SDK 02:Google genai SDK——Gemini 与 Vertex AI 双后端及 Function Calling 进阶

在 2024 年底之前,Google 的大模型开发生态一直让工程师相当困扰:面向个人开发者的 Google AI Studio 使用的是旧版 google-generativeai 库,而面向企业级生产环境的 Vertex AI 则使用完全不同的 google-cloud-aiplatform 库。两套 SDK 的包名不同、对象结构不同、函数调用协议也不同,项目从本地 PoC 走向云端生产往往需要重构大量适配代码。

Google 推出的全新统一 Python SDK google-genai 彻底解决了这一割裂。它不仅统一了 Gemini 模型在个人端与云端的调用规范,更在架构设计上展现了一个关键特性:通过 Vertex AI 后端,这一套 SDK 可以直接调用包括 Claude、Llama 3.1、Mistral 在内的第三方头部模型。


安装与初始化架构

全新统一 SDK 安装:

1
pip install google-genai

如果需要使用企业级 Vertex AI 后端,需配合安装认证库并完成 GCP 登录:

1
2
pip install google-cloud-aiplatform
gcloud auth application-default login

双后端设计模式(Dual Backend Pattern)

google-genai 最核心的架构设计是**“一套接口,两套后端”**。业务层所有模型调用的代码完全保持一致,仅仅在客户端实例化时切分流量目的地:

方式一:Google AI Studio(个人开发与快速验证)

直接读取 API 密钥,仅支持访问 Gemini 自身系列模型:

1
2
3
4
5
import os
from google import genai

# 使用个人 API Key 初始化
client = genai.Client(api_key=os.environ["GOOGLE_API_KEY"])

方式二:Vertex AI(企业级生产与多模型中心)

传入 GCP 项目与区域配置,不仅可访问 Gemini,还可直接调度托管在 Model Garden 中的第三方模型:

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
from google import genai

# 使用 Vertex AI 企业后端初始化
client = genai.Client(
vertexai=True,
project="my-gcp-enterprise-project",
location="us-central1",
)

# 统一 API 调用 Claude 3.5 Sonnet
claude_response = client.models.generate_content(
model="claude-3-5-sonnet-v2@20241022",
contents="用简明的技术语言解释一下分布式事务的 2PC 机制。",
)
print(claude_response.text)

# 统一 API 调用开源 Llama 3.1
llama_response = client.models.generate_content(
model="meta/llama-3.1-405b-instruct-maas",
contents="Explain raft consensus algorithm briefly.",
)
print(llama_response.text)

业务逻辑代码完全不需要因切换模型供应商而更换 SDK。


对话与参数配置

基础生成与系统指令配置

在基础生成中,系统指令(System Instruction)与推理参数封装在 types.GenerateContentConfig 中:

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
import os
from google import genai
from google.genai import types

client = genai.Client(api_key=os.environ["GOOGLE_API_KEY"])

response = client.models.generate_content(
model="gemini-2.0-flash",
contents="def calculate_tax(amount): return amount * 0.1 有什么潜在隐患?",
config=types.GenerateContentConfig(
system_instruction="你是一个金融级支付系统架构师,专注于浮点数精度与边界异常审查。",
temperature=0.1,
max_output_tokens=1024,
),
)
print(response.text)

多轮对话(Chat Session)

SDK 提供了 chats 对象管理上下文历史,免去手动维护消息数组的成本:

1
2
3
4
5
6
7
8
chat = client.chats.create(model="gemini-2.0-flash")

# 第一轮
r1 = chat.send_message("我们目前系统有 3 个微服务节点,日订单量约 50 万单。")
# 第二轮(模型自动携带历史上下文)
r2 = chat.send_message("基于刚才的体量,数据库读写分离是否是当务之急?")

print(r2.text)

Function Calling 的两套实现模式

在实现工具调用时,google-genai 兼顾了“白盒细粒度控制”与“黑盒开箱即用”两种诉求。

模式一:手动状态循环(细粒度控制)

当需要对工具执行增加严格的权限审查、审计日志、或在调用前插入人工审批时,应当采用手动循环模式。

工具调用包含三步标准时序:

  1. 声明工具 Schema:向模型提供函数签名;
  2. 捕获 function_call:模型判定需要调用工具并返回参数;
  3. 回送 function_response:本地执行完函数后,将结构化产物灌回上下文。

代码实现

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
from google import genai
from google.genai import types


def query_db_node_status(node_id: str) -> dict:
"""查询指定数据库节点的运行状态"""
db = {
"db-node-primary": {"role": "master", "status": "UP", "lag_ms": 0},
"db-node-replica-1": {"role": "slave", "status": "UP", "lag_ms": 45},
}
return db.get(node_id, {"error": "节点不存在"})


# 1. 显式构建工具定义
tools_declaration = types.Tool(
function_declarations=[
types.FunctionDeclaration(
name="query_db_node_status",
description="查询数据库集群中节点的角色、健康状态与复制延迟",
parameters=types.Schema(
type=types.Type.OBJECT,
properties={
"node_id": types.Schema(
type=types.Type.STRING,
description="节点 ID,如 'db-node-primary'",
)
},
required=["node_id"],
),
)
]
)

tools_map = {"query_db_node_status": query_db_node_status}

# 2. 手动驱动状态机
contents = [
types.Content(
role="user",
parts=[types.Part(text="检查 db-node-replica-1 节点的复制延迟情况。")],
)
]

while True:
response = client.models.generate_content(
model="gemini-2.0-flash",
contents=contents,
config=types.GenerateContentConfig(tools=[tools_declaration]),
)

first_part = response.candidates[0].content.parts[0]

# 检查是否触发了工具调用
if hasattr(first_part, "function_call") and first_part.function_call:
fc = first_part.function_call
print(f"[拦截并执行工具] 函数: {fc.name}, 参数: {dict(fc.args)}")

# 本地真实执行
fn_output = tools_map[fc.name](**dict(fc.args))

# 将模型的调用意图作为 model 消息追加
contents.append(
types.Content(
role="model",
parts=[types.Part(function_call=fc)],
)
)

# 将执行结果作为 user 消息追加回上下文
contents.append(
types.Content(
role="user",
parts=[
types.Part(
function_response=types.FunctionResponse(
name=fc.name,
response={"result": fn_output},
)
)
],
)
)
else:
# 模型已输出最终文本
print(f"\n[最终结论]\n{response.text}")
break

模式二:自动模式(Automatic Function Calling)

如果业务对工具调用过程没有额外的审计或拦截需求,直接将普通 Python 函数传入 tools 列表,并开启 automatic_function_calling 即可:

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
def query_db_node_status(node_id: str) -> str:
"""查询指定数据库节点的运行状态与复制延迟。"""
return f"节点 {node_id}: 状态正常, 复制延迟 12ms"


# 直接传函数,SDK 自动完成 Schema 提取与内部轮转
response = client.models.generate_content(
model="gemini-2.0-flash",
contents="帮我看看 db-node-replica-1 是否正常?",
config=types.GenerateContentConfig(
tools=[query_db_node_status],
automatic_function_calling=types.AutomaticFunctionCallingConfig(disable=False),
),
)

print(response.text)

原生多模态处理

Gemini 在多模态上的能力覆盖非常完整,支持直接在 contents 中混排文本、图片、PDF、音频乃至视频文件:

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
from google.genai import types

# 读取本地图片或架构图
with open("architecture_diagram.png", "rb") as f:
img_data = f.read()

response = client.models.generate_content(
model="gemini-2.0-flash",
contents=[
types.Part(
inline_data=types.Blob(
mime_type="image/png",
data=img_data,
)
),
types.Part(text="请分析图中展现的架构拓扑,指出其单点故障(SPOF)风险。"),
],
)
print(response.text)

对于数百兆的音视频与长文档,推荐配合 SDK 内置的 client.files.upload 接口上传至云端缓存后再行调用。


常用模型选型参考

模型 ID 特点与核心优势 推荐应用场景
gemini-2.0-flash 极低时延与成本,支持原生流式与多模态 日常生产主力,高并发 Agent 循环与工具调用
gemini-2.0-flash-thinking 原生思维链(CoT)推理,逻辑严密 复杂算法审查、系统架构权衡推演
gemini-1.5-pro 具备 1M ~ 2M 超长上下文吞吐能力 全库代码分析、大型合同/研报整本理解

工程总结

  1. 统一 API 解耦基础设施:优先使用 google-genai。开发期通过 api_key 使用 AI Studio 低成本验证,生产期通过 vertexai=True 切换到企业专网与多模型池,逻辑零修改;
  2. 按场景选择 Function Calling 模式:生产核心风控与写操作采用手动循环以保障白盒可控;内部只读查询任务采用自动模式以精简代码;
  3. 压榨多模态与长上下文红利:在处理长篇幅文档、音视频分析等复杂多媒体任务时,Gemini 原生格式具备显著工程优势。

下一篇我们将拆解 Anthropic 官方 SDK,看看以严谨著称的 Claude 是如何通过 Messages API 贯彻低层透明控制哲学的。


系列导航与参考