Sharing Knowledge: Scientific Communication - SSOAR

Weitere Magazine

Empfehlungen

Info

162 Judith Plümer The summarizers generate SOIF records for each document. SOIF was designed for the storage and exchange of the summaries of documents which may originate in different formats. It was built as a part of the harvest software. The structure of SOIF is simple and effective. A SOIF object consists of multiple attribute-value pairs. For example author{16}: Erwin Mustermann is such a pair, the number in brackets gives the number of characters of the value. There are no restrictions on the number and names of attributes. This principle of attribute-value pairs makes SOIF equivalent to the use of the HTML2.0 META tag with its NAME-CONTENT pairs. But it can also be used to store information from the body of a HTML document or from other formats. For example, the HTML summarizer of harvest assigns all words between and to the header attribute. Obviously, it makes sense to build such attribute-value pairs from any markup language like HTML, SGML, XML or TeX/LaTeX. The summarizer for documents in a typeset format as .dvi, .ps or .pdf work differently, they only store the first 100 words of these documents as keywords because the text is not semantically structured. One could think of criteria to extract more structure out of these unstructured documents via methods of artificial intelligence, heuristics or the use of thesauri (see below). The summarizers of harvest also store metadata such as the mimetype of the original document, its URL, URLs of links contained in the document, name and host of the gathering agent, time of generation, update time and time to live. All these pieces of information play an important role for the system because the information is not constant. They are used to control the expiration times. Collection of documents and problems of heterogeneity Since SOIF objects are just ASCII files, they can be transferred by any protocol as SOIF streams. Harvest builds a database or an index of SOIF objects by using Glimpse which makes the attribute-value pairs searchable. Now the problem arises when upgrading to HTML4.0 or to XML with metadata in RDF (Resource Description Framework). In MPRESS we currently store SCHEME information of HTML4.0 META as a qualifier. For example is mapped to DC.subject.msc{5}: 19D10
MPRESS - transition of metadata formats 163 This is done by a modification of harvest’s HTML summarizer. But now under Glimpse you won’t get this object when searching for DC.subject=19D10. This does not create any problems in a controlled environment, but when you start to import SOIF documents from remote sites your metadata become heterogeneous. This problem currently occurs in MPRESS with the metadata files we get from the arXiv.org mirror at Augsburg: The name of the author of a paper is stored there as DC.creator and not as DC.creator.PersonalName as inside of MPRESS. This problem of heterogeneity is solved during the import of the SOIF documents into the broker. The gatherer command that manages this import pipes the SOIF objects through a Perl script that substitutes the older DC coding of the author’s name with the newer version. In analogue to the SCHEME problem the handling of sub qualifiers in DC coded in RDF is difficult at this stage. After collecting the documents and generating SOIF objects for them, each gathering agent exports the SOIF objects via a daemon. Now the broker component collects the SOIF objects from several daemons and computes an index of these preprints, or more exactly an index of their SOIF objects. The gatherers that feed MPRESS are also distributed over the world which makes an upgrading of software or used formats difficult. The user of MPRESS can only see the broker which is an interface for searching in this information pool. Not all stored metadata elements can be used and evaluated by the user. Currently searches for authors, titles and keywords (and dates in France) are supported, but if you know the correct syntax of the other metadata items, you can also use the full text search field to query for the other items. An important feature of MPRESS is the browse function besides the search functionality. Here the user gets a table for the several levels of the MSC and the number of preprints classified in the respective categories. Then he can browse through the next level of the MSC or search in the respective subset of preprints. To achieve such a functionality of the broker, one has to refine the perl scripts of the harvest distribution which organize the preparation of the query to the glimpse database and which arrange the answers from the glimpse database to be nicely displayed. In detail this is the BrokerQuery.pl script in harvest1.4 or the nph-search script in harvest1.5. Especially the browse functionality is not part of the harvest software. The pages that appear during the browsing process are generated pseudo-automatically: The pages reside in the file system of the WWW server, but they are generated once a day to be up to date. This is done by a cronjob, which is a perl script that queries the broker for the different MSC categories. Analogously the MSC specific search interfaces are generated on the fly to send an MSC enriched query to the broker that filters out the papers without the correct MSC codes. That means the browsing function is
Seite 1:
Sharing Knowledge: Scientific Commu
Seite 4 und 5:
Tagungsberichte Herausgegeben vom I
Seite 6 und 7:
Die Deutsche Bibliothek - CIP-Einhe
Seite 8 und 9:
6 Inhalt Infrastrukturen für innov
Seite 11:
Vorwort Zur neunten Frühjahrstagun
Seite 14 und 15:
12 Heike Andermann ted, in 1994-199
Seite 16 und 17:
14 Heike Andermann schaftlerInnen e
Seite 18 und 19:
16 Heike Andermann tung beibehalten
Seite 20 und 21:
18 Heike Andermann NBII). 26 Für d
Seite 23 und 24:
Qualitätssicherung und Nutzung von
Seite 25 und 26:
Seite 27 und 28:
Seite 29 und 30:
Seite 31 und 32:
Seite 33 und 34:
Seite 35 und 36:
Seite 37 und 38:
Seite 39 und 40:
vascoda Das gemeinsame Portal von I
Seite 41 und 42:
vascoda - Das gemeinsame Portal von
Seite 43 und 44:
Seite 45 und 46:
Seite 47:
Seite 50 und 51:
48 Klaus Hahn Abstract The advancem
Seite 52 und 53:
50 Klaus Hahn her“ [6]). So wäre
Seite 54 und 55:
52 Klaus Hahn men auch als Begriff
Seite 56 und 57:
54 Klaus Hahn Arbeiten notwendig (>
Seite 58 und 59:
56 Klaus Hahn Fazit Zur effektiven
Seite 61 und 62:
Unterstützung kooperativer Verfahr
Seite 63 und 64:
Seite 65 und 66:
Seite 67 und 68:
Seite 69 und 70:
Seite 71 und 72:
Seite 73 und 74:
PhysNet und seine Spiegel - Das Pro
Seite 75 und 76:
Seite 77 und 78:
Seite 79 und 80:
Seite 81 und 82:
Seite 83 und 84:
Seite 85 und 86:
Online-Hochschulschriften für die
Seite 87 und 88:
Seite 89 und 90:
Seite 91 und 92:
Seite 93 und 94:
Seite 95:
Seite 98 und 99:
96 Rudi Schmiede, Stephan Körnig s
Seite 100 und 101:
98 Rudi Schmiede, Stephan Körnig s
Seite 102 und 103:
100 Rudi Schmiede, Stephan Körnig
Seite 104 und 105:
Seite 106 und 107:
Seite 108 und 109:
Seite 110 und 111:
108 Jutta von Maurice strument lief
Seite 112 und 113:
110 Jutta von Maurice bare Kopien v
Seite 114 und 115: 112 Jutta von Maurice einschließli
Seite 116 und 117: 114 Jutta von Maurice hebungen. Dem
Seite 118 und 119: 116 Jutta von Maurice http://www.df
Seite 121 und 122: Maßnahmen zur Förderung der Infor
Seite 133 und 134: LIMES - A System for a Distributed
Seite 139: LIMES - A System for a Distributed
Seite 142 und 143: 140 Frank Oldenettel, Michael Malac
Seite 160 und 161: 158 Judith Plümer prints werden du
Seite 162 und 163: 160 Judith Plümer modification, th
Seite 166 und 167: 164 Judith Plümer based on a few p
Seite 168 und 169: 166 Judith Plümer Migration from H
Seite 170 und 171: 168 Judith Plümer Contact Judith P
Seite 172 und 173: 170 Dennis Reil für traditionelle
Seite 174 und 175: 172 Dennis Reil gartner (2002) ist
Seite 176 und 177: 174 Dennis Reil können. Wichtig f
Seite 178 und 179: 176 Dennis Reil nahezu identisch, s
Seite 180 und 181: 178 Dennis Reil Insgesamt kann also
Seite 183 und 184: Reflections on the Value Chain of S
Seite 193: Reflections on the Value Chain of S
Seite 196 und 197: 194 Natascha Schumann, Wolfgang Mei
Seite 207 und 208: ViFaPhys - Virtuelle Fachbibliothek
Seite 215 und 216:
Weiterentwicklung von digitalen Bib
Seite 217 und 218:
Seite 219 und 220:
Seite 221 und 222:
Seite 223 und 224:
Seite 225 und 226:
Seite 227:
Seite 230 und 231:
228 Markus Kalb, Günther Specht ve
Seite 232 und 233:
230 Markus Kalb, Günther Specht Di
Seite 234 und 235:
232 Markus Kalb, Günther Specht Da
Seite 236 und 237:
234 Markus Kalb, Günther Specht 4.
Seite 238 und 239:
236 Markus Kalb, Günther Specht pe
Seite 240 und 241:
238 Markus Kalb, Günther Specht [O
Seite 242 und 243:
240 Maximilian Stempfhuber � Info
Seite 244 und 245:
242 Maximilian Stempfhuber Weiterle
Seite 246 und 247:
244 Maximilian Stempfhuber Zum zwei
Seite 248 und 249:
246 Maximilian Stempfhuber Abbildun
Seite 251 und 252:
Das didaktische Metadatensystem DML
Seite 253 und 254:
Seite 255 und 256:
Seite 257 und 258:
Seite 259 und 260:
Seite 261 und 262:
Seite 263 und 264:
Seite 265 und 266:
Seite 267 und 268:
Seite 269 und 270:
The C 2 M project: a wrapper genera
Seite 271 und 272:
Seite 273 und 274:
Seite 275 und 276:
Seite 277 und 278:
Seite 279 und 280:
Seite 281 und 282:
Seite 283 und 284:
Seite 285 und 286:
Analyse der Qualität der multimedi
Seite 287 und 288:
Seite 289 und 290:
Seite 291 und 292:
Seite 293 und 294:
Seite 296:
Ziele des erstmals von der Initiati
Alle anzeigen

Sharing Knowledge: Scientific Communication - SSOAR

Sie wollen auch ein ePaper? Erhöhen Sie die Reichweite Ihrer Titel.

Template löschen?

Als Template speichern?