Home
/
Topics
/
Document retrieval

Topic

Document retrieval

About: Document retrieval is a research topic. Over the lifetime, 6821 publications have been published within this topic receiving 214383 citations.

...read moreread less

Papers published on a yearly basis

2023
2022
2021
2020
2019
2018
2017
2016
2015
2014
2013
2012
2011
2010
2009
2008
2007
2006
2005
2004
2003
2002
2001
2000
1999
1998
1997
1996
1995
1994
1993
1992
1991
1990
1989
1988
1987
1986
1985
1984
1983
1982
1981
1980
1979
1978
1977
1976
1975
1974
1973
1972
1971
1970
1969

1 / 2

Papers

PDF

Open Access

More filters

Journal Article•DOI•

Ranking documents with a thesaurus.

[...]

Roy Rada¹, Ellen J. Bicknell²•Institutions (2)

University of Liverpool¹, National Institutes of Health²

01 Sep 1989-Journal of the Association for Information Science and Technology

TL;DR: This article reports on exploratory experiments in evaluating and improving a thesaurus through studying its effect on retrieval, and how adding non-BT relations to MeSH showed how these non- BT relations could improve document ranking, if DISTANCE were also appropriately revised to treat these relations differently from BT relations.

...read moreread less

Abstract: This article reports on exploratory experiments in evaluating and improving a thesaurus through studying its effect on retrieval. A formula called DISTANCE was developed to measure the conceptual distance between queries and documents encoded as sets of thesaurus terms. DISTANCE references MeSH (Medical Subject Headings) and assesses the degree of match between a MeSH-encoded query and document. The performance of DISTANCE on MeSH is compared to the performance of people in the assessment of conceptual distance between queries and documents, and is found to simulate with surprising accuracy the human performance. The power of the computer simulation stems both from the tendency of people to rely heavily on broader-than (BT) relations in making decisions about conceptual distance and from the thousands of accurate BT relations in MeSH. One source for discrepancy between the algorithms' measurement of closeness between query and document and people's measurement of closeness between query and document is occasional inconsistency in the BT relations. Our experiments with adding non-BT relations to MeSH showed how these non-BT non-BT relations to MeSH showed how these non-BT relations could improve document ranking, if DISTANCE were also appropriately revised to treat these relations differently from BT relations.

...read moreread less

93 citations

Patent•

Document retrieval using fuzzy-logic inference

[...]

Kousuke Takahashi¹, Karvel K. Thornber¹•Institutions (1)

NEC¹

15 Aug 1994

TL;DR: In this article, the results of a full-text, document search by a character string search processor are treated as vector patterns whose elements become a term match grade by use of a membership function of the term match frequency.

...read moreread less

Abstract: The results of a full-text, document search by a character string search processor are treated as vector patterns whose elements become a term match grade by use of a membership function of the term match frequency. The closest pattern to the query pattern is found by the similarity between the query pattern and each of the filed sample patterns. The similarity is calculated by use of fuzzy-logic. The similarity is ranked in order of similarity magnitude, thereby reducing the search time. The search time can be shortened by categorizing the filed patterns by term set and similarity to a cluster center pattern. If the cluster center patterns are stored, the closest cluster address can be inferred by fuzzy logic inference from the match between the query document and the term set or the similarity of the query to the cluster center.

...read moreread less

93 citations

Patent•

Document retrieval system with access control

[...]

Steven T. Kirsch¹•Institutions (1)

Google¹

10 Sep 1997

TL;DR: In this paper, an electonic document retrieval system and method for a collection of information distributed over a network having documents stored in web or document servers in which an access control list relates user identification to documents to which a user has access.

...read moreread less

Abstract: An electonic document retrieval system and method for a collection of information distributed over a network having documents stored in web or document servers in which an access control list relates user identification to documents to which a user has access. No access control lists are contained in the documents themselves nor are comparisons made between lists of users, with their access levels, and the classifications of documents. Rather, by the use of URLs or pointers, it is possible to associate every document to which a user has access with the user identification number or code. URLs have a hierchical format which allows partial URLs to indicate levels of access. HTTP protocol, FTP and CGI protocol employ URL calls for documents and can use the access control method and system of the present invention. When a search query is applied to a query server, a list of hits is returned, together with pertinent URLs. The query server consults each access control list associated with each document server, to present to the user only those URLs for which he has a proper access level. Other URLs for which the user does not have proper access are kept hidden from the user.

...read moreread less

93 citations

Proceedings Article•

Spoken document retrieval for TREC-7 at Cambridge University

[...]

S.E. Johnson, Pierre Jourlin, G.L. Moore, Karen Sparck Jones, Philip C. Woodland - Show less +1 more

01 Jan 1998

92 citations

Proceedings Article•DOI•

Chinese text retrieval without using a dictionary

[...]

Aitao Chen, Jianzhang He, Liangjie Xu, Fredric C. Gey¹, Jason Meggs¹ - Show less +1 more•Institutions (1)

University of California, Berkeley¹

01 Jul 1997

TL;DR: The results show that, for all three sets of queries, the simple bigram indexing and the purely statistical word segmentation perform better than the popular dictionary-based maximum matching method with a dictionary of 138,955 entries.

...read moreread less

Abstract: It is generafly believed that words, rather than characters, should be the smallest indexing unit for Chinese text retrieval systems, and that it is essential to have a comprehensive Chinese dictionary or lexicon for Chhmse text retrieval systems to do well. Chinese text has no delimiters to mark woni boundaries. As a result, any text retrieval systems that build word-based indexes need to segment text into words. We implemented several statistical and dictionary-hazed word segmentation methods to study the effect on retrieval effectiveness of different segmentation methods using the TREC-S Chinese test collection and topics. The results show that, for all three sets of queries, the simple bigram indexing and the purely statistical word segmentation perform better than the popular dictionary-based maximum matching method with a dictionary of 138,955 entries.

...read moreread less

92 citations

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
…
80
81
82
83
84
85
86
…
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200

Collapse

Network Information

Performance

Metrics

6,866

Papers

224,605

Citations

No. of papers in the topic in previous years
Year	Papers
2023	9
2022	39
2021	107
2020	130
2019	144
2018	111

Document retrieval

Papers published on a yearly basis

Papers

Trending Questions (10)

Network Information

Related Topics (5)

Performance

Metrics