Home
/
Topics
/
Speaker recognition

Topic

Speaker recognition

About: Speaker recognition is a research topic. Over the lifetime, 14990 publications have been published within this topic receiving 310061 citations.

...read moreread less

Papers published on a yearly basis

2023
2022
2021
2020
2019
2018
2017
2016
2015
2014
2013
2012
2011
2010
2009
2008
2007
2006
2005
2004
2003
2002
2001
2000
1999
1998
1997
1996
1995
1994
1993
1992
1991
1990
1989
1988
1987
1986
1985
1984
1983
1982
1981
1980
1979
1978
1977
1976
1975
1974
1973
1972
1971
1970
1969

1 / 2

Papers

PDF

Open Access

More filters

Patent•

Handwritten and voice control of vehicle components

[...]

Gabriel Ilan, Benjamin Giloh, Arie Kadosh

06 May 1999

TL;DR: A recognition system for use in a vehicle or the like includes a handwriting recognizer (44) and a voice recognizer(38) for receiving handwriting and voice signals, where the signals are associated with commands used to operate with a variety of vehicle appliances.

...read moreread less

Abstract: A recognition system (20) for use in a vehicle or the like includes a handwriting recognizer (44) and a voice recognizer (38) for receiving handwriting and voice signals, where the signals are associated with commands used to operate with a variety of vehicle appliances. Such appliances may include, but are not limited to, car alarms (32), electric windows, personal computers (28), navigation systems (26), and audio (30) and telecommunications equipment (24).

...read moreread less

68 citations

Proceedings Article•DOI•

Accent classification in speech

[...]

S. Deshpande¹, Sharat Chikkerur¹, Venu Govindaraju¹•Institutions (1)

University at Buffalo¹

17 Oct 2005-Scopus

TL;DR: This paper distinguishes between standard American English and Indian Accented English using the second and third formant frequencies of specific accent markers to achieve a suitable classification for these two accent groups.

...read moreread less

Abstract: Apart form the word content and identity of a speaker; speech also conveys information about several soft biometric traits such as accent and gender. Accurate classification of these features can have a direct impact on present speech systems. An accent specific dictionary or word models can be used to improve accuracy of speech recognition systems. Gender and accent information can also be used to improve the performance of speaker recognition systems. In this paper, we distinguish between standard American English and Indian Accented English using the second and third formant frequencies of specific accent markers. A GMM classification is used on the feature set for each accent group. The results show that using just the formant frequencies of these accent markers is sufficient to achieve a suitable classification for these two accent groups.

...read moreread less

68 citations

Proceedings Article•DOI•

Speaker change detection and speaker clustering using VQ distortion for broadcast news speech recognition

[...]

K. Mori¹, Seiichi Nakagawa¹•Institutions (1)

Toyohashi University of Technology¹

07 May 2001

TL;DR: The aim is to apply speaker grouping information to speaker adaptation for speech recognition by using vector quantization (VQ) distortion as the criterion and showing the superiority of the proposed method.

...read moreread less

Abstract: Addresses the problem of the detection of speaker changes and clustering speakers when no information is available regarding speaker classes or even the total number of classes. We assume that no previous information on speakers is available (no speaker model, no training phase) and that people do not speak simultaneously. The aim is to apply speaker grouping information to speaker adaptation for speech recognition. We use vector quantization (VQ) distortion as the criterion. A speaker model is created from successive utterances as a codebook by a VQ algorithm, and the VQ distortion is calculated between the model and an utterance. A result was obtained by the experiment on speaker detection and speaker clustering. The speaker change detection experiment was compared with results by generalized likelihood ratio and Bayesian information criterion. We show the superiority of our proposed method.

...read moreread less

68 citations

Proceedings Article•DOI•

Multimodal speaker detection using error feedback dynamic Bayesian networks

[...]

Vladimir Pavlovic, Ashutosh Garg¹, James M. Rehg², Thomas S. Huang¹•Institutions (2)

University of Illinois at Urbana–Champaign¹, Hewlett-Packard²

01 Jan 2000

TL;DR: This work forms a learning framework for DBNs based on error-feedback and statistical boosting theory and applies this framework to the problem of audio/visual speaker detection in an interactive kiosk environment using "off-the-shelf" visual and audio sensors.

...read moreread less

Abstract: Design and development of novel human-computer interfaces poses a challenging problem: actions and intentions of users have to be inferred from sequences of noisy and ambiguous multi-sensory data such as video and sound. Temporal fusion of multiple sensors has been efficiently formulated using dynamic Bayesian networks (DBNs) which allows the power of statistical inference and learning to be combined with contextual knowledge of the problem. Unfortunately simple learning methods can cause such appealing models to fail when the data exhibits complex behavior. We formulate a learning framework for DBNs based on error-feedback and statistical boosting theory. We apply this framework to the problem of audio/visual speaker detection in an interactive kiosk environment using "off-the-shelf" visual and audio sensors (face, skin, texture, mouth motion, and silence detectors). Detection results obtained in this setup demonstrate superiority of our learning framework over that of the classical ML learning in DBNs.

...read moreread less

68 citations

Proceedings Article•DOI•

Speaker identification by combining MFCC and phase information in noisy environments

[...]

Longbiao Wang¹, Kazue Minami², Kazumasa Yamamoto², Seiichi Nakagawa²•Institutions (2)

Shizuoka University¹, Toyohashi University of Technology²

14 Mar 2010

TL;DR: The effectiveness of phase information for noisy environments on speaker identification in noisy environments with integrated MFCC with phase information is described.

...read moreread less

Abstract: In conventional speaker recognition methods based on MFCC, the phase information has been ignored. Recently, we proposed a method that integrated MFCC with the phase information on a speaker recognition method. Using the phase information, the speaker identification error rate was reduced by 78% for clean speech. In this paper, we describe the effectiveness of phase information for noisy environments on speaker identification. Integrationg MFCC with phase information, the speaker error identification rates were reduced by 20%∼70% in comparison with using only MFCC in noisy environments.

...read moreread less

68 citations

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
…
182
183
184
185
186
187
188
…
189
190
191
192
193
194
195
196
197
198
199
200

Collapse

Network Information

Performance

Metrics

15,632

Papers

337,766

Citations

No. of papers in the topic in previous years
Year	Papers
2023	165
2022	468
2021	283
2020	475
2019	484
2018	420

Speaker recognition

Papers published on a yearly basis

Papers

Trending Questions (10)

Network Information

Related Topics (5)

Performance

Metrics