feat(agent): add voice message support with TTS/STT for Telegram and WeChat

- Integrate voice message handling: detect and extract audio references from Telegram and WeChat messages, route to agent with voice reply preference.
- Add voice provider abstraction and OpenAI-based TTS/STT implementation.
- Implement agent tool `send_voice_message` for generating and sending voice replies, with fallback to text if voice is unavailable.
- Extend agent prompt and context to support voice reply instructions.
- Update notification and message schemas to support audio fields.
- Add Telegram and WeChat voice sending logic, including audio file conversion and temporary media upload for WeChat.
- Add tests for voice helper and agent voice routing.
This commit is contained in:
jxxghp
2026-04-12 12:30:02 +08:00
parent 9dababbcfd
commit e5f97cd299
17 changed files with 945 additions and 167 deletions
+6
View File
@@ -55,6 +55,8 @@ class CommingMessage(BaseModel):
callback_query: Optional[Dict] = None
# 图片列表(图片URL或file_id
images: Optional[List[str]] = None
# 语音/音频引用列表
audio_refs: Optional[List[str]] = None
def to_dict(self):
"""
@@ -86,6 +88,10 @@ class Notification(BaseModel):
text: Optional[str] = None
# 图片
image: Optional[str] = None
# 语音文件路径
voice_path: Optional[str] = None
# 语音消息附带说明文字
voice_caption: Optional[str] = None
# 链接
link: Optional[str] = None
# 用户ID