Tren
dar
대시보드
논문
뉴스
GitHub
AI 소식
KO
EN
로그인
AI 소식
논문
대시보드
뉴스
GitHub
“speech”
이 키워드와 관련된 논문 · GitHub · 뉴스를 한곳에 모았습니다.
논문
12
전체 →
Semantic Scholar
음성·오디오
인용 355
확장 가능한 스트리밍 음성 합성을 위한 대규모 언어 모델 기반 CosyVoice 2
CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models
OpenAlex
ML 방법론
인용 148
순차 데이터 생성을 위한 잠재 변수 순환 신경망
A Recurrent Latent Variable Model for Sequential Data
arXiv
음성·오디오
인용 0
풀더플스피치언어모델의 음향-의미 간섭 원인 규명 및 계층적 분리 모델 제안
Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs
arXiv
음성·오디오
인용 0
단어 수준 음향 특성을 명시적으로 제어하는 LLM 기반 음성 합성 프레임워크
WordVoice: Explicit and Decoupled Multi-Dimensional Word-Level Control for LLM-Based TTS
arXiv
안전·보안
인용 0
AI 안전 평가를 위한 언어적 모호성과 명령 충돌 벤치마크 제안
Adversarial Pragmatics for AI Safety Evaluation: A Benchmark for Instruction Conflict, Embedded Commands, and Policy Ambiguity
arXiv
ML 방법론
인용 0
저지연 온디바이스 AI 서빙을 위한 실행 상태 캡슐과 그래프 기반 체크포인트/복원 메커니즘
Execution-State Capsules: Graph-Bound Execution-State Checkpoint and Restore for Low-Latency, Small-Batch, On-Device Physical-AI Serving
arXiv
생성모델
인용 0
자연어 장면 묘사로 야생 오디오 환경의 다중 화자 대화 생성
Reference-Driven Multi-Speaker Audio Scene Generation from In-the-Wild Priors
arXiv
음성·오디오
인용 0
트랜스듀서를 결합해 LLM 음성인식의 실시간 스트리밍을 가능하게 한 방법
TRADE: Transducer-Augmented Decoder for Speech LLM
bioRxiv
자연어·LLM
인용 0
딥 앙상블 기반 뇌-컴퓨터 인터페이스의 실시간 음성 디코딩 성능 검증
Neural decoding of speech using deep neural ensembles
bioRxiv
자연어·LLM
인용 0
인간 뇌의 음운-의미 순차 변환 메커니즘이 AI 음성 이해를 향상시킨다
Human-like sequential sound-to-meaning transfer drives artificial speech comprehension
Semantic Scholar
자연어·LLM
인용 985
생성형 AI로 환자-의사 대화를 자동으로 SOAP/BIRP 진료 기록으로 변환
Intelligent Clinical Documentation: Harnessing Generative AI for Patient-Centric Clinical Note Generation
Semantic Scholar
자연어·LLM
인용 4
LLM의 구조화된 출력 품질을 다중 소스로 평가하는 벤치마크
The Structured Output Benchmark: A Multi-Source Benchmark for Evaluating Structured Output Quality in Large Language Models
GitHub
5
전체 →
음성·오디오
Python
★ 30.5K
토크나이저 없는 다국어 TTS, 음성 디자인 및 클로닝 지원
OpenBMB/VoxCPM
제품·출시
Python
★ 11.9K
오픈소스 모델로 로컬 보이스 에이전트를 구축하는 모듈형 파이프라인
huggingface/speech-to-speech
음성·오디오
Python
★ 18.8K
170배 빠른 산업용 음성인식 툴킷, 화자 분리·감정 탐지·스트리밍 지원
modelscope/FunASR
음성·오디오
뉴스
5
전체 →
Hacker News
제품·출시
▲ 4
ALS 환자, 뇌파로 말하고 풀타임 근무 가능해져
AI and brain-computer interface allow speechless ALS patient to work full-time
Google DeepMind
제품·출시
▲ 0
구글, 70개 언어 실시간 음성 번역 모델 Gemini 3.5 Live Translate 출시
Fluid, natural voice translation with Gemini 3.5 Live Translate
Hacker News
Python
★ 3.6K
고충실도·고표현력 음성 및 음향 생성 오픈소스 모델군
OpenMOSS/MOSS-TTS
에이전트
Python
★ 4.5K
오픈소스 음성 AI 에이전트 플랫폼
dograh-hq/dograh
제품·출시
▲ 4
AI 에이전트에 이미지·영상·음성 생성 기능을 통합 제공하는 오픈소스 프레임워크
Show HN: Imagent – agentic image/video/speech generation
Hacker News
▲ 13
If AI Outputs Aren't Speech, Who Has to Prove They're Human?
OpenAI
▲ 0
How we built a realtime system for responsive voice AI in six months