Monday, 22 August 2011

Costs & Benefits

A key question arose throughout the project, not least in the two archivists' focus groups - is Linked Data worth the input of professional staff time? From the front end perspective, are the improvements for users - enriched catalogues published more speedily and improved, automated, linking with external services - sufficient to justify the extra effort required from staff? Aren't Google, Autonomy/HP and other large corporations that manage huge quantites of data doing enough already, or will do very soon (a Google 'Linked Data' button, anyone?). A fundamental point is that archivists are under enormous pressure to justify and quantify potential benefits of Linked Data to senior management through the simplication of often confusing and obscure terminology and by the use of exemplars and online test areas.

The OMP project showed that initial professional scepticism can be overcome if Linked Data can be simply defined and the benefits clearly set out. Archivists will use Linked Data if a service or services are provided that automate of simplify mark-up or the semantic process more generally and embed it within existing cataloguing workstreams. Ideally, these can be built out of trusted aggregations, authorities or cataloguing systems such as the Hub, AIM25, TNA, CALM or ATOM. They are less likely to use Linked Data if it is perceived to be a complex, though potentially useful, add-on requiring detailed specialist knowledge and delivered without support or guidance ('built it and they will come'). The ability to retroconvert legacy catalogues and CLDs with Linked Data through automation against OpenCalais and other engines will help sell Linked Data more effectively, as can validation of metadata created out of mass digitisations and OCR.

The OMP project has underlined the value of Linked Data in a number of ways:

  • Increased access and discovery
  • Increased use and return on investment in cataloguing (speeding up cataloguing, enabling tools that require an archivist to locate and link information - for example indexing, finding already-existing authority records and linking to them; finding suitable subject terms; locating places from geonames or similar)
  • Enhanced ability to justify expenditure on services and resource development (improved web-hits and connecting with heavily used services)
  • Exposure of information to novel and different uses (Combining ALM collections for the delivery of services, including commercial services - apps, exhibitions, mapping, new tools etc)
The specific benefits, as demonstrated via AIM25 are:

Updated workflow interface including:
  • Reduction of the requirement for archivists to input HTML
  • Reduction of the on-screen size of the form
  • Integration of the process of selecting access points
  • Automatic semantic annotation to aid selection of classifying terms
  • Authority lookup (internal and external - UKAT, GeoNames, etc) to improve rigour of metadata
Semantic rendering of the classification terms used by AIM25 (separate from the AIM25 access-points records):
  • SKOS representation of AIM25-UKAT data
  • RDF for AIM25 people, families and corporate names
  • GeoNames representation of AIM25 place data
Use of RDFa where available to enhance the public interface of AIM25
  • Semantic lookup allowing users to further explore definitions and instances of terms based on the properties defined during the workflow process.

The main business case is two-fold: adding value and boosting efficiency. Archivists are very attracted to the idea of enabling UKAT in Linked Data but as an active service like OpenCalais, not a look-up. AIM25 has developed a SKOS version of UKAT and a workflow tool that would link from a revised AIM25 data entry template to a LD UKAT.

Of place, personal name, corporate name and subject, subject terms are arguably the most subjective, requiring the archivist to exercise judgment on the preferred term with the collection and potential users in mind. OMP has shown that subject terms throw up the least accurate semantic returns from a linguistic analysis service such as OpenCalais (places can often be matched with absolute precision, as can personal names). OMP has improved professional efficiency by developing a hover tool to enable the archivist to select a preferred subject term from UKAT or via connecting to LD versions of LCSH/NRA and to add this term or terms to their new catalogue/CLD.

Without such automation, Linked Data won't be embedded or the data linked will be limited in scope. Flexibility is key. Focus group archivists concluded that they need the ability to analyse as much or as little of a description as they need, and to reach that faceting decision as speedily as possible - selecting the most important entities that require linking in any body of text, and fields (just 'creator', 'institution' etc or terms within Scope and Content or Admin/Biographical?). The value of broader authority data was reiterated by the archivists - analysis should not be limited to Scope and Content. A fundamental point is that back-end Linked Data enhancement works best when it works with the grain of professional practice - pragmatically and speedily.

The OMP approach is innovative in that it offers further exposure of data - and all AIM25 data has been processed as part of the project. Sustainability will be maintained going forward either by periodic manual data dumps into OpenCalais or by automated calendared refreshes - the same approach could be envisaged for LD UKAT as a national service plugged into local systems such as CALM. Improving the OpenCalais vocabularly by importing archive-specific terms is crucial to the success of mark-up. Analysis of the catalogue data is only valuable if OpenCalais learns from archivists. Until this happens, the breadth of vocabulary will limit the scope of the mark-up. It is also worth putting pressure on the main suppliers of archival cataloguing software to encourage them to embed support in periodic upgrades.

Experimentation with NRA data is ongoing - this will test how difficult it would be to build an authorities service off the NRA/ARCHON. The results will be described in a separate blog post.


Thursday, 28 July 2011

Licensing

Linked data, amongst the many challenges it presents, requires licensing which is appropriate to its intended uses. As the compilation of databases is not regarded as a creative act under at least US law, the Creative Commons licence is probably not appropriate for licensing linked data. Instead, the Open Data Commons licences  (http://opendatacommons.org/licenses/) defined by the Open Knowledge  Foundation appears a more appropriate choice for this purpose.

Open Data Commons includes three licences: the Public Domain Dedication and License (PDDL), which places the data in the public domain and waives all rights, the Attribution License (ODC-By) which allows the sharing and adaptation of the data provided it remains attributed, and the Open Database License (ODbL), which allows the same rights provided any adaptations are distributed under the same licence.

Links to these licences are provides as RDF triples on the website: for  instance:-
·                 rdf:RDF
·                 xmlns:cc='http://creativecommons.org/ns#'
·                 xmlns:rdf='http://www.w3.org/1999/02/22-rdf-syntax-ns#'
·                 xmlns:dcq='http://purl.org/dc/terms/'
·                 cc:License rdf:about="http://opendatacommons.org/licenses/odbl/1.0/">
·                 cc:legalcode
·                 rdf:resource="http://opendatacommons.org/licenses/odbl/1.0/"/>
·                 dcq:hasVersion>1.0</dcq:hasVersion>
·                 cc:License>
·                 rdf:RDF>

to define the Open Database License. They can therefore be readily incorporated into any linked data application.  Discussions with partner institutions within AIM25 will tease out the preferred licence. It is expected to be ODbL.

As a matter of record a statement is being added to all descriptions reflecting AIM25 as the origin of the data, that AIM25 is a partnership in which contributing members have rights, but that the data is otherwise freely accessible for reuse as it has been since inception some 11 years ago. (see for example the quantities of AIM25 created entries for the Women’s Library now also offered in the Hub.) The statement will reference the use of a Open Data Commons licence. 

An issue still to be addressed however is thesaurus support. UKAT is based on UNESCO which gave permission for development but we assume further permission will be required for further use and embedding.  The same applies to MESH, and Gay and Lesbian and other vocabularies developed by partner organisation.  We have also used Getty Arts and Architecture selectively. 

Gareth Knight and Patricia Methven

Wednesday, 27 July 2011

Second Archivists' Focus Group

A second archivists' focus group was convened on Wednesday 27th July. In attendance were representatives of Senate House Library, Wandsworth Heritage Services, London School of Economics, ULCC, King's College London, London Metropolitan Archives, the Institute of Education and the National Archives.

The group reviewed progress since the last meeting and viewed a demo of the new back end editing area of AIM25. The back end area allows individual, several or all ISAD(G) fields in collection level descriptions to be analysed against the existing AIM25 version of UKAT, a fuzzy match option against UKAT to identify synonyms and against OpenCalais using the OpenCalais service. These returns are listed alphabetically alongside the record in collapsable lists and these entities also highlighted in the text using a colour-coding formula to distinguish subject, place name, personal name and corporate names, and triples where those were identified. Analysis normally takes a matter of seconds though timing-out can occur with longer records.

A mouse roll-over feature has been inserted for each term in the text, allowing users to identify the particular attribute of the term in a tick box drop down menu, using a connect to one or more external services including Geonames (is 'London' the capital of the UK, a place in Canada, an author or part of a corporate name?). Triples can also be interrogated in roll-overs in order that the editing archivist might validate or clarify these entities. These choices can be saved and then exported. The enhanced content will be expressed in a new front-end delivery for the test records that demonstrate linking with external services, in order to enhance the user experience by pulling together reliable external information on a place, name, subject etc relevant to that collection.

The debate centred on how the editing process can be speeded up - for example by 'signing-off' the capital-of-the-UK version of London for all examples of 'London' across all ISAD(G) fields after review of the first instance ('treat all subsequent examples of 'London' in this record as the capital'). Linking with the NRA is desirable to identify authority terms set out in NCA format. Linking with Library of Congress was raised as an important deliverable in order to maximise the opportunity for synergy between archive and library descriptions, particularly in local authority record offices.

The question of updating UKAT in an ongoing fashion was raised - maintaining an RDF version of AIM25 UKAT must require minimum ongoing effort given constrained budgets and workloads. Analysis reveals the limitations of the existing thesaurus but also the possibility and desirability for external services like OpenCalais to be enhanced by input from ALM thesauri and vocabularies. This requires a conversation between JISC or key UK institutions and OpenCalais and similar services.

Next steps are to improve and tidy the editing area (for example by changing colour coding), plugging this into the front-end for test records and exploring NRA/ARCHON collaboration.

Monday, 18 July 2011

Archivists' focus group

King's College Archives recently hosted a focus group comprising leading London archivists familiar with using AIM25. The purpose of the focus group was to understand how Linked Data approaches might speed up the behind-the-scenes editing work of the archivist and improve the front-end user experience. Representatives of Senate House Library, the London School of Economics, Wandsworth Heritage Services, the London Metropolitan Archives, the British Postal Museum, the Institute of Education and University of London Computer Centre were in attendance. Development work on new AIM25 records was showcased.

Real-time use of OpenCalais was demonstrated and tested by members using sample data and the results compared. Subject-term creation was shown to be an area of potential concern - OpenCalais was developed by Reuters as a news and current affairs-support service and terms tend to reflect this focus. More input from archive vocabularies was called for to enable OpenCalais's corpus to be enriched with Higher Education and other terminology. It was also suggested that Linked Data could provide fuzzy matching between formal if rather arcane UNESCO-style subject terms and terms that are in more popular use, to encourage discovery and take-up. It was suggested that the UK Archival Thesaurus could be enhanced and made available in a SKOS version.

The practical use to hard-pressed archivists came up time and again as a topic of conversation. Most archivists have neither the time nor budgets to engage in experimentation but need practical tools that they can plug into their work without fuss. Quantifying the benefits of Linked Data is vital to sell the approach to funders and institutional management. Cross-domain services are an important attraction in surfacing and linking archive information with books and museum content. The benefits of linking to Wikipedia services (DPedia) were raised - Wikipedia lies at the centre of the Linked Data universe. Biographical content could be imported wholesale from other sources and adapted for use in a particular record, which would save time researching and writing one from scratch.

The plans of proprietary suppliers like Axiell and Adlib was raised as an issue - are they planning to incoporate Linked Data tools in future versions of their archive management software? The role of Google was discussed. Do they have any Linked Data plans and if not, why not?

The issue was raised of which fields in ISAD(G) to include in Linked Data work. It was argued that focusing only on Scope and Content was a mistake, not least because of the value of authority records (Admin/Biographical) and related records fields. Linking to the NRA to surface related collections was discussed.

Indexing was discussed by panel members. Editors of AIM25, the Archives Hub or similar tools should be able to draw on Linked Data to improve or enhance the personal, corporate and place names of new and existing records (and the ability to retrospectively run existing records through OpenCalais was flagged as an important requirement - archivists are more likely to embrace LD if they can painlessly re-index their current content). Linked Data provides the opportunity for more automation  and speedier indexing, which are particularly useful for smaller archives without cataloguing expertise.

Next steps included further development on the indexing tools in order to compare workflow with traditional methods; build a prototype front-end delivery system to enhance collection level descriptions and engage in conversation with Google and others to identify best practice.

Tuesday, 10 May 2011

URIs for AIM25 access metadata

We've been giving some thought to a suitable URI scheme to adopt for this project which could mesh with the requirements of the current AIM25 metadata requirements.

The current AIM25 system contains four types of access records: personal names (Person), corporate names (Organisation), subject (Subject) and geographic names (Place). To allow semantic linkages to be formed, we shall need a coherent set of URIs that can handle all of these.

For Subjects, one option may be the UKAT thesaurus which is available as SKOS RDF. Each concept here has already been assigned a URI: for instance, for 'Poetry' this is http://www.ukat.org.uk/thesaurus/concept/525. Note that UKAT is no longer being edited, however, and so may become out-of-date in the future.

For place names, there are several gazeteers available: the Archives Hub recommends the Getty Thesaurus (http://www.getty.edu/research/conducting_research/vocabularies/tgn/index.html),
but there is not yet a set of URIs authorised by the Getty (although they are working on this).

The LOCAH (Linked Open Copac Archives Hub - http://blogs.ukoln.ac.uk/locah/) projectproduced a set of guidelines for URIs which look very useful. For each of these categories they would take this form:-


Personal and corporate name names

LOCAH recommends the following format for personal names:-

{root}/id/person/{rules}/{person-name}

so, once we have decided on a suitable root - let's say for the moment http://data.aim25.ac.uk/, for Burns | John | 1774-1868 | surgeon, we'd have

http://data.aim25.ac.uk/id/person/ncarules/burnsjohn1774-1868surgeon.

Similarly for corporate names, we'd have:-

http://data.aim25.ac.uk/concept/organisation/nra/stthomasshospitallondon

For geographic names, we'd have:-

http://data.aim25.ac.uk/id/place/ncarules/grimsbylincolnshireta2709


and finally for subjects:-

http://data.aim25.ac.uk/id/concept/aim25subjects/medicalsciences


These seem to be viable options, although a final decision has yet to be reached.

Friday, 6 May 2011

A machine readable layer for AIM25

I've been busy on the AIM25 test server adding a machine readable layer.

Of course AIM25 has for a long time offered the metadata held for each collection as EAD. I've taken this as my jumping off point and had a go at adding some more formats to the AIM25 arsenal that will hopefully be of user to any silicon based users of the AIM25 service.

There are still a few screws to tighten but hopefully this work will represent a useful tool for the OMP work to enrich the AIM25 metadata.

---

First off I wanted to mimic the browsing structure that us humans take for granted as we make our way from the homepage to collection page on the website. For this we used Encoded Archival Context (EAC) to list and describe the Institutions and their Collections.

Next we wanted to extend the work started by Gareth and Richard using semantic web services at the collection level. Once at this level we can access EAD (as we always could), to this we added the dynamically generated output from OpenCalais (OCS) and a modified version of EAD with the OCS output embedded in the content (EAD aft. OCS). This latter is also dynamically generated.

Lastly we added some browser-side scripting to the original HTML pages to highlight terms identified by openCalais. All of the above uses openCalais dynamically so be patient. Obviously the goal would be to use a triple-store generated using OC at the point of creation (and change) of records.

This work so far is really a demonstration of some possible ways of expressing and enriching AIM25 content. It is by no means an exhaustive (or even authoritative) list of possible formats, but we hope it will serve to make tangible some of the ideas we've been discussing over the past month or so.

Thanks to our firewall-wallahs you can now browse AIM25-OMP here (thankfully in HTML too).

If nothing else this has been a good exercise in getting to know AIM25 a bit better and whipping my XSLT into a useful shape and of course dipping my toe in the semantic ocean.

Thursday, 24 March 2011

Semantic Analysis of AIM25 EAD


Rory and I met with Richard Gartner and Gareth Knight at CeRch today, to catch up with their investigations into using GATE and OpenCalais to process the EAD outputs from AIM25.

Results look very encouraging. OpenCalais, in particular, generates a post-processing set of identified entities (personal names, place names, corporate names) which Richard G has then created regular expressions to locate these in the body of the EAD and wrap in appropriate EAD tags (<persname> etc).

This suggests that the way forward for enhancing the existing data entry processes for AIM25 will involve dispatching the EAD-compliant data entered by collections manager to OpenCalais, and returning the data, with enhanced markup, for checking by the submitter. This hook should be easy enough to insert for manual, form-based entry; for batch entry processes we will need to assess whether any significant delays are introduced.

We've also started to consider ideas for a URI scheme for the entities identified. Our current working hypothesis is that this will involve defining a "data" namespace for AIM25, binding to http://data.aim25.ac.uk/. Within that we can develop a structure along the lines /person, /place, /corporate_body, and append our unique IDs for each entity. Further research is necessary, particularly into the recommendations of the Cabinet Office recommendations for Designing URI Sets for the UK Public Sector.

These URIs can then be used in identifier attributes for our EAD elements (<persname>, etc.), and thence easily transformed into an RDFa format for the Web-based HTML rendering of the AIM25 catalogues.

Next steps include further investigating how to implement and assert relationships between our entities and other open datasets (e.g. our_entity  is_the_same_as  your_entity). And how to make the authority data, duly marked-up, available as open metadata.

Rory and I can now start to consider suitable approaches to embedding this in our development copy of the existing AIM25 system, and we'll continue to liaise closely with CeRch for advice on  the relative merits of Gate and OpenCalais processing, and guidance on URI implementation.