Home
/
Authors
/
Jakob Uszkoreit

Author

Jakob Uszkoreit

Other affiliations: University of California, Berkeley

Bio: Jakob Uszkoreit is an academic researcher from Google. The author has contributed to research in topics: Machine translation & Transformer (machine learning model). The author has an hindex of 36, co-authored 84 publications receiving 37432 citations. Previous affiliations of Jakob Uszkoreit include University of California, Berkeley.

Papers published on a yearly basis

2022
2021
2020
2019
2018
2017
2016
2015
2013
2012
2011
2010
2009
2008

Papers

PDF

Open Access

More filters

Patent•

Generating neural network outputs using insertion operations

[...]

Jakob Uszkoreit¹, Mitchell Stern¹, Jamie Kiros, William Chan•Institutions (1)

Google¹

23 Jul 2020

Patent•

Paramètres d'optimisation pour la traduction automatique

[...]

Wolfgang Macherey, Shankar Kumar, Roy W. Tromble, Franz Josef Och, Jakob Uszkoreit - Show less +1 more

02 Jul 2009

TL;DR: In this paper, a mise en œuvre de l'invention fournit un procede comprend l'acces a un espace d'hypothese; the realisation du decodage de l'sespace d'Hypothese pour obtenir une hypothese de traduction reduisant au minimum the classification d'une erreur attendue, calculee par rapport a un space de justifications; and l'apport de l"hypothess de traductions obtenue que l'util

...read moreread less

Abstract: L'invention concerne des procedes, des systemes et un appareil, y compris des produits-programmes d'ordinateur, pour la traduction de langues Une mise en œuvre de l'invention fournit un procede Ce procede comprend l'acces a un espace d'hypothese; la realisation du decodage de l'espace d'hypothese pour obtenir une hypothese de traduction reduisant au minimum la classification d'une erreur attendue, calculee par rapport a un espace de justifications; et l'apport de l'hypothese de traduction obtenue que l'utilisateur pourra utiliser en tant que traduction suggeree dans une traduction cible

...read moreread less

Patent•

Optimieren von parametern für maschinenübersetzung

[...]

Wolfgang Macherey, Shankar Kumar, Roy W. Tromble, Franz Josef Och, Jakob Uszkoreit, Ignacio Thayer - Show less +2 more

02 Jul 2009

Patent•

어텐션-기반의 시퀀스 변환 신경망

[...]

Noam Shazeer, Aidan N. Gomez, Lukasz Kaiser, Jakob Uszkoreit, Llion Jones, Niki Parmar, Illia Polosukhin, Ashish Vaswani - Show less +4 more

31 Jul 2019

TL;DR: In this article, the authors proposed a method to solve the problem of the lack of resources in the South Korean market by using the concept of "social media" and "social networks".

...read moreread less

Abstract: 방법들, 시스템들 및 디바이스들은 입력 시퀀스로부터 출력 시퀀스를 생성하기 위한 컴퓨터 저장 매체 상에 인코딩된 컴퓨터 프로그램을 포함한다. 일 양태에서, 시스템들 중 하나는 입력 시퀀스를 수신하고 네트워크 입력의 인코딩된 표현을 생성하도록 구성된 인코더 신경망과, 상기 인코더 신경망은 하나 이상의 인코더 서브네트워크의 시퀀스를 포함하고, 각 인코더 서브네트워크는 입력 위치들 각각에 대한 개별 인코더 서브네트워크 입력을 수신하여 입력 위치들 각각에 대한 개별 서브네트워크 출력을 생성하도록 구성되고, 그리고 각 인코더 서브네트워크는 입력 위치들 각각에 대한 서브네트워크 입력을 수신하여, 상기 입력 순서의 각각의 특정 입력 위치에 대해, 상기 특정 입력 위치에서 상기 인코더 서브네트워크 입력으로부터 도출된 하나 이상의 쿼리를 사용하여 상기 인코더 서브네트워크 입력들에 어텐션 메커니즘을 적용하는 인코더 셀프-어텐션 서브-계층을 포함한다.

...read moreread less

Patent•

Query composition system

[...]

Jakob Uszkoreit¹•Institutions (1)

Google¹

28 Sep 2015

TL;DR: In this paper, the authors describe a system for generating data describing context clusters and context cluster probabilities, where each context cluster includes query inputs based on the input context for each of the query inputs and the content described by each query input, and each cluster probability indicates a probability that at a query input that belongs to the context cluster will be selected by the user, receiving, from a user device, an indication of a user event that includes data indicating a context of the user device.

...read moreread less

Abstract: Methods, systems, and apparatus for generating data describing context clusters and context cluster probabilities, wherein each context cluster includes query inputs based on the input context for each of the query inputs and the content described by each query input, and each context cluster probability indicates a probability that at a query input that belongs to the context cluster will be selected by the user, receiving, from a user device, an indication of a user event that includes data indicating a context of the user device, selecting as a selected context cluster, based on the context cluster probabilities for each of the context clusters and the context of the user device, a context cluster for selection input by the user device, and providing, to the user device, data that causes the user device to display a context cluster selection input that indicates the selected context cluster for user selection.

...read moreread less

1
2
3
4
5
6
7
8
9
10
11
12
13
…
14
15
16
17

Collapse

Cited by

PDF

Open Access

More filters

Posted Content•

BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

[...]

Jacob Devlin¹, Ming-Wei Chang¹, Kenton Lee¹, Kristina Toutanova¹•Institutions (1)

Google¹

11 Oct 2018-arXiv: Computation and Language

TL;DR: A new language representation model, BERT, designed to pre-train deep bidirectional representations from unlabeled text by jointly conditioning on both left and right context in all layers, which can be fine-tuned with just one additional output layer to create state-of-the-art models for a wide range of tasks.

...read moreread less

Abstract: We introduce a new language representation model called BERT, which stands for Bidirectional Encoder Representations from Transformers. Unlike recent language representation models, BERT is designed to pre-train deep bidirectional representations from unlabeled text by jointly conditioning on both left and right context in all layers. As a result, the pre-trained BERT model can be fine-tuned with just one additional output layer to create state-of-the-art models for a wide range of tasks, such as question answering and language inference, without substantial task-specific architecture modifications. BERT is conceptually simple and empirically powerful. It obtains new state-of-the-art results on eleven natural language processing tasks, including pushing the GLUE score to 80.5% (7.7% point absolute improvement), MultiNLI accuracy to 86.7% (4.6% absolute improvement), SQuAD v1.1 question answering Test F1 to 93.2 (1.5 point absolute improvement) and SQuAD v2.0 Test F1 to 83.1 (5.1 point absolute improvement).

...read moreread less

29,480 citations

Proceedings Article•DOI•

BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

[...]

Jacob Devlin¹, Ming-Wei Chang¹, Kenton Lee¹, Kristina Toutanova¹•Institutions (1)

Google¹

11 Oct 2018

TL;DR: BERT as mentioned in this paper pre-trains deep bidirectional representations from unlabeled text by jointly conditioning on both left and right context in all layers, which can be fine-tuned with just one additional output layer to create state-of-the-art models for a wide range of tasks.

...read moreread less

Abstract: We introduce a new language representation model called BERT, which stands for Bidirectional Encoder Representations from Transformers. Unlike recent language representation models (Peters et al., 2018a; Radford et al., 2018), BERT is designed to pre-train deep bidirectional representations from unlabeled text by jointly conditioning on both left and right context in all layers. As a result, the pre-trained BERT model can be fine-tuned with just one additional output layer to create state-of-the-art models for a wide range of tasks, such as question answering and language inference, without substantial task-specific architecture modifications. BERT is conceptually simple and empirically powerful. It obtains new state-of-the-art results on eleven natural language processing tasks, including pushing the GLUE score to 80.5 (7.7 point absolute improvement), MultiNLI accuracy to 86.7% (4.6% absolute improvement), SQuAD v1.1 question answering Test F1 to 93.2 (1.5 point absolute improvement) and SQuAD v2.0 Test F1 to 83.1 (5.1 point absolute improvement).

...read moreread less

24,672 citations

Journal Article•DOI•

Squeeze-and-Excitation Networks

[...]

Jie Hu¹, Li Shen², Samuel Albanie², Gang Sun¹, Enhua Wu¹ - Show less +1 more•Institutions (2)

Chinese Academy of Sciences¹, University of Oxford²

18 Jun 2018

TL;DR: This work proposes a novel architectural unit, which is term the "Squeeze-and-Excitation" (SE) block, that adaptively recalibrates channel-wise feature responses by explicitly modelling interdependencies between channels and finds that SE blocks produce significant performance improvements for existing state-of-the-art deep architectures at minimal additional computational cost.

...read moreread less

Abstract: The central building block of convolutional neural networks (CNNs) is the convolution operator, which enables networks to construct informative features by fusing both spatial and channel-wise information within local receptive fields at each layer. A broad range of prior research has investigated the spatial component of this relationship, seeking to strengthen the representational power of a CNN by enhancing the quality of spatial encodings throughout its feature hierarchy. In this work, we focus instead on the channel relationship and propose a novel architectural unit, which we term the “Squeeze-and-Excitation” (SE) block, that adaptively recalibrates channel-wise feature responses by explicitly modelling interdependencies between channels. We show that these blocks can be stacked together to form SENet architectures that generalise extremely effectively across different datasets. We further demonstrate that SE blocks bring significant improvements in performance for existing state-of-the-art CNNs at slight additional computational cost. Squeeze-and-Excitation Networks formed the foundation of our ILSVRC 2017 classification submission which won first place and reduced the top-5 error to 2.251 percent, surpassing the winning entry of 2016 by a relative improvement of ${\sim }$ ∼ 25 percent. Models and code are available at https://github.com/hujie-frank/SENet .

...read moreread less

14,807 citations

Posted Content•

RoBERTa: A Robustly Optimized BERT Pretraining Approach

[...]

Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Michael Lewis, Luke Zettlemoyer, Veselin Stoyanov - Show less +6 more

26 Jul 2019-arXiv: Computation and Language

TL;DR: It is found that BERT was significantly undertrained, and can match or exceed the performance of every model published after it, and the best model achieves state-of-the-art results on GLUE, RACE and SQuAD.

...read moreread less

Abstract: Language model pretraining has led to significant performance gains but careful comparison between different approaches is challenging. Training is computationally expensive, often done on private datasets of different sizes, and, as we will show, hyperparameter choices have significant impact on the final results. We present a replication study of BERT pretraining (Devlin et al., 2019) that carefully measures the impact of many key hyperparameters and training data size. We find that BERT was significantly undertrained, and can match or exceed the performance of every model published after it. Our best model achieves state-of-the-art results on GLUE, RACE and SQuAD. These results highlight the importance of previously overlooked design choices, and raise questions about the source of recently reported improvements. We release our models and code.

...read moreread less

13,994 citations

Posted Content•

An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

[...]

Alexey Dosovitskiy¹, Lucas Beyer¹, Alexander Kolesnikov¹, Dirk Weissenborn², Xiaohua Zhai¹, Thomas Unterthiner¹, Mostafa Dehghani¹, Matthias Minderer¹, Georg Heigold², Sylvain Gelly¹, Jakob Uszkoreit¹, Neil Houlsby¹ - Show less +8 more•Institutions (2)

Google¹, German Research Centre for Artificial Intelligence²

22 Oct 2020-arXiv: Computer Vision and Pattern Recognition

TL;DR: Vision Transformer (ViT) attains excellent results compared to state-of-the-art convolutional networks while requiring substantially fewer computational resources to train.

...read moreread less

Abstract: While the Transformer architecture has become the de-facto standard for natural language processing tasks, its applications to computer vision remain limited. In vision, attention is either applied in conjunction with convolutional networks, or used to replace certain components of convolutional networks while keeping their overall structure in place. We show that this reliance on CNNs is not necessary and a pure transformer applied directly to sequences of image patches can perform very well on image classification tasks. When pre-trained on large amounts of data and transferred to multiple mid-sized or small image recognition benchmarks (ImageNet, CIFAR-100, VTAB, etc.), Vision Transformer (ViT) attains excellent results compared to state-of-the-art convolutional networks while requiring substantially fewer computational resources to train.

...read moreread less

12,690 citations

1
2
3
4
…
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200

Collapse