Friday, 3 August 2012

Getting Closer to a Map Interface

We're continuing to get closer to a useful geographical interface for visualizing collections.  Using the AIM25-CALM service that Rory's created, we're able to call relevant collections based on time and latitude/longitude filters.  From the programming side, Rory created a test interface to experiment with what kind of data we would get with various calls to the service (fig 1).

[Fig 1]. Early tests of map showing collections returned from AIM25-CALM service.
One of the things that we've been experimenting with is the granularity, which defines how many levels of detail exist for a placename.  For instance: country/province/city/ or country/state/county/city/. These administrative districts vary by country, but helps us lower the signal to noise ratio. As you can see in Figure 2, the level 4 granularity returns more specific information in France than it does in the United Kingdom.

[Fig 2]. Experimenting with maximum granularity yields more results in France than the UK.  
As we move forward with this, we'll be further testing the returns of the service, and refining the user interface to integrate into the Historypin environment.

Wednesday, 4 July 2012

CALM User Testing

We organised a focus group session on Thursday 28th June at the Wellcome Trust to review progress so far on adapting the CALM backend to query external services and generate and store RDF. The meeting comprised a mix of CALM archivists and some professionals familiar with cataloguing processes. The main purpose of the meeting was to use a CALM 9.3 development environment to test the robustness of workflows for analysing catalogue and authority data; comment on the quality and sources of external data; and review improvements to the front end - CALMVIEW - that will publish appropriate service links. The CALMVIEW linking between test 9.3 installation and internet could not be configured on the day, as CALMVIEW development is still under way, but a screenshot of the AIM25 link was shown to the participants.

The underlying rationale is that Linked Data processing, sharing and exporting from CALM should become as normal and integral part of cataloguing as is possible - one that does not require an immense investment in additional work or process on the part of archivists, or a detailed, and unrealistically obtainable, technical knowledge of RDF.

The main workflow for analysis of sample Wellcome Library and Cumbria catalogues was tested out using the UKAT service, DBpedia and the BL British National Bibliography, to cover archival, biographical and bibliographical-type material. Key improvements/findings requested were:

  • Improved bulk analysis of records (to speed up processing)
  • Preview of the resolved multiple service returns before embedding (to overcome the problem of poor quality external data being selected or similar-sounding names of people and places being mistakenly chosen, or, for example, to preview and select the correct edition of a multi-edition printed publication)
  • The ability for archivists in the back end to refine and select only certain records for publication (necessary because some services only return dirty data or data strings, which are of little value to researchers)
  • Demarcation of front end presentation of Linked Data links from host catalogue data to minimise confusion as to the origin of the data source  
  • The need for the archive profession to agree to the creation of a priority list of Linked Data services that would be useful for professionals and users, such as the NRA and specialist vocabularies.
The focus group was followed by a meeting of the CALM User Group, at which CALM representatives outlined the release schedule.

Tuesday, 3 July 2012

User Interface and First Steps Toward Meshups

When we originally conceptualized how we might begin to automatically incorporate collections data into Historypin, we imagined having the ability to peer into the past and be able to reach into various collections nearby to pull out relevant information.  Ultimately, this is still what we're working toward, but a number of complications prevent this from being feasible at the moment.  But it's worth sharing some of our mockups for how we saw this working.

[Fig 1] This mockup shows the existing experience of historical photos overlaid in Street View on the Historypin site.  We've added a "Dig Deeper" panel on the right which initiates a call to the Step change service based on the date and location.
[Fig 2] This mockup shows the results of the query to the Step change service, including  relevant AIM25 collections that may relate to the date and location referenced.
[Fig 3]  Selecting one of the collections would return information more information about the collection, including the location of the collection and a link to the collection webpage.
There are a number of reasons why this execution is not quite practical yet, though it may be feasible in a future project.  The primary complication here is the signal to noise ratio when in London, as so many of the collections within AIM25 are relevant to London, but the geographic specificity in the collection metadata is often not very detailed on the level of granularity that you get when in Street View.  If we ask for relevant collections within a 1 mile radius of a specific latitude and longitude for instance, we may get back 2-300 collections with little clue as to why this collection is relevant to this location. 

Another problem is what we've been calling the Needle In A Haystack problem, once you get away from London and into other parts of the world.  While the AIM25 collections are largely in Greater London (sorry if I'm getting my terminology wrong--I'm an American!), there are many collections that are relevant to other parts of the world.  Rory has done an amazing job parsing out locations from the collections metadata and using Geonames to resolve these locations.  So we can now see that a particular collection may have relevance to locations in China for instance, which is one of the locations  we've been testing with.  Here, our problem is that we've got just one or two collections and they are geotagged for a small town where someone lived.  So unless we set a really large bounding box, unless you happen to be in Street View in that town, you'd never learn about that collection, even though it has documents pertinent to many locations in China and Tibet.

Friday, 8 June 2012

Development of CALM 9.3: towards a simple linked data indexing tool

One of the key components of our Step change project is the development of Axiell's software products CALM and CALMVIEW to support some linked data functionality. Both applications are widely used within the UK archival community. The current public release of CALM is version 9.2, and I have been supplied with a release of 9.3 to use with our catalogue data in a test environment here.

There has been for some years a structure of keen and active CALM user groups and a formal dialogue between these and Axiell in regard to ongoing refinement and development of these products. Very simply, within the time and resource confines of the Step change project, we are looking to develop the following tools that will:
  1. allow CALM users to interrogate linked data services, return results and insert selected URIs into CALM records
  2. allow these links to be displayed in the CALMVIEW web front end to the CALM application
  3. allow CALMVIEW to expose data from the CALM application in RDF so that it in turn can be interrogated by other services
In this entry I will be discussing development of point 1 (points 2 and 3 will emerge a little later in the project). Basically there are two aspects to the functionality in v.9.3 i) administration facility to allow configuration/testing of current and new linked data services and ii) user's facility to link URIs into CALM records. Let's look at the administration part first:
From the main administration menu, we go to a Linked Data submenu and can see the default linked data services which have been configured. These are AIM25, British Library British National Bibliography (BNB) and Wikipedia (Dbpedia). At this point we can add, remove or edit/test these services.

We can also see how the link to AIM25 is initially configured:

Next we can see the XSLT used to transform XML received from the selected service into CALM's XML.

At this point we can also select the test function and send some text to search the service and see what is returned. Here we have searched the AIM25 service with the text 'Churchill' and can see some XML returned.

This is transformed by the XSLT to CALM XML

and we can see the results processed below:

You can also use the Admin menu to determine which databases in CALM can be potentially linked to that service - you are not restricted to your catalogue database for example. So that's the admin part. What about linking a CALM catalogue record to a URL from one of these services? Well in catalogue menu, you select the Authorities menu in the left, and you now see a Linked Data button.  


Now you decide which service to link to:
You now decide which fields you would like the service to search against, and you can add your own free text to generate extra searches. Here I've selected the title field which contains the text 'Clementine Churchill', which is what I want to search against AIM25:


Now I see the results coming back from AIM25, and I've selected the Clementine Churchill person authority
And I then use the utility to post this URL into the CALM catalogue record
Here I've done the same with the Dbpedia service - you can see I get two Clementine Churchill topic returns. These actually turn out to be duplicates after checking the URLs so I end up having to delete one from the catalogue record...
Here's the catalogue record with links to both services in
Testing is all in the very early stages but some interesting (and I think important) points have emerged already:

1. Configuring or adding new Linked Data services to CALM: Axiell have designed this process to be generic and extensible. However to do this requires a knowledge and ability to write XSLT which will be beyond most archivists (me included!). So where does this leave us? A couple of options will likely emerge. Firstly, new services may be added in CALM upgrades through the User Group requests for enhancements process. Secondly, those CALM users with access to technical assistance may well be able to produce the requisite XSLT transforms necessary to add new services and share these through the User Group process.

 2. Making CALM find the internet: The Linked Data function in CALM means that the application now needs to reach the internet to do its searching. This is potentially problematic and will very much depend on the particular user technical environment and how user access to the internet is configured. Axiell, myself and the technical support for Cumbria County Council ICT have spent some considerable time testing and configuring to make this happen. Here in Cumbria for use a proxy server to access the internet and users have to enter username and password credentials when prompted. The proxy server then compares these credentials against a database, and if correct, allows access to the internet. Axiell have therefore had to develop CALM to prompt for these internet credentials if required, and we have found this works. In the near future we are likely to upgrade to a new proxy server which will check the user's Active Directory credentials and work straight off these, providing a cleaner, simpler route to the internet, and this should bypass the need to enter your internet credentials in CALM. However until CALM 9.3 is in a range of other user environments, it will not be possible to ensure that internet access is ensured.

3. Finding and determining meaningful results: Even after just a few minutes' use one can see some interesting issues emerge. The BL BNB service doesn't usually return anything from the text found in the fields of most of our CALM catalogue records, as it interprets text within each field as a complete search term. However using the freetext extra search box (without any CALM fields selected) to enter a crisper more precise word or string of text usually does work. The Dpbedia results seen in CALM from a search certainly need to be treated with caution. For example when parsing a catalogue record relating to some local police documents with fields which contained various text such as 'Whitehaven Police Station', 'Crime', 'Law', I had results back from Dpbedia as follows: 'Whitehaven', 'Crime', 'Law'. The Whitehaven URL related in fact to an entry about Whitehaven Railway Station, the Crime URL related not to an entry about criminalilty but to an entry for a Californian rock band called Crime, and the Law URL related not to legislation but to an entry about the actor Jude Law. Hmmmm....all good food for thought about future work in this area.

4. Useful services?: As things stand, three default services were configured because they were readily accessible and of some possible use to archivists and archive users. But the key services for archivists are as yet still in development or not readily accessible (I'm thinking of things like National Register of Archives, Manorial Documents Register etc). And there are major issues of course relating to person/corporate name authorities or geo name authorities which are essential areas for development, possibly by way of some sort of national brokering service to make these available in the same way that AIM25 is providing a sort of de facto UK Archival Thesaurus (UKAT) brokering service.

I'll be spending the next few weeks using CALM 9.3 to provide links to many hundreds of catalogue records, so we can see the results in a test version of CALMVIEW and which will allow proper user testing.

Thursday, 7 June 2012

Historypin and Step change: Collections in Context

The not-for-profit behavior change agency We Are What We Do is known for creating simple yet compelling tools that encourage people to change their behavior in small ways that amount to big impacts in areas like waste reduction, childhood obesity, social isolation, etc.


Historypin is a project we created together with Google and several cultural memory institutions to help bridge the gap between cultures and generations and help rebuild and strengthen ties within communities. We created simple tools that could be used by individuals, schools, communities, and institutions to create a shared view of the layers of history that make up a community.

It's been nearly a year since the official launch of Historypin and we've experienced tremendous growth, with hundreds of cultural heritage institutions adding content and creating Channels, and tens of thousands of individual users joining the site and adding their own content.  It's in this growth that we've started to realize the potential for academic research of the blending of multiple collections, and the opportunity to incorporate other sources of content in more automated ways.  We've just begun to scratch the surface of what will be possible in years to come.

Our role in the Step change project is to explore options for incorporating AIM25 collection holdings into the geographic landscape of scores of historical photos and audio/video recordings; and to assist researchers, scholars, and the general public in the discovery and contextualization of these holdings.

In the coming weeks, we'll explore some of the user interface prototypes we're testing and document the challenges and successes we meet along the way.

Monday, 16 April 2012

Problems being addressed

The main purpose of Step change is to find a way for archive institutions to adopt Linked Data approaches when they otherwise face constraints of time, financial and personnel resources. Quite simply, Linked Data is unlikely to be adopted in the archives community if it is considered a potentially expensive add-on when apparently higher priorities can be identified such as cataloguing backlogs, problems with the continuity of digital data, the need for improved catalogues and search engine optimisation to make even the most basic descriptions available to the public for the first time.

Archivists need to be able to demonstrate to senior management improved productivity from using Linked Data approaches. Step change will try and do this by developing a workflow tool to enable archivists to analyse catalogues against various external services such as Geonames, and to capture the resulting RDF. The tool is being designed in such a way as to make cataloguing per se easier to carry out, for example by allowing consistent and accurate indexing through interrogation of UKAT. In this way, the creation of RDF will become a normal part of catalogue creation and management. The roll-out to CALM version 9.3 is designed to embed this approach across the CALM customer base.

The second problem is how to link archive catalogues effectively with other types of content such as bibliographic records, museum records, external databases and maps. The challenges here are considerable: the lack of availability of viable external services and APIs within the cultural sector; licensing issues with the reuse of map data out of contenxt; the lack of an available historical gazetteer to map antique placenames; discontinuities in describing name authorities; and the absence of user testing to determine which sources ought to be linked and at which levels of granularity. How much data is too much data for users to assimilate? The project is seeking to address this problem of usability by testing the mixing of several external data sets including the National Register of Archives, Historypin and Wikipedia and then user testing with CALM users and with local authority users in the north west of England. The key here is showing that publication of useful data side by side about a person, place or theme adds to the user experience, speeds up research or opens up new avenues of research.

Future work could include the blending of research datasets, the mapping of previously unmaped data and the use of the Workflow tool to crowdsource RDF creation using existing datasets, for example image metadata. These would all help support a stronger use-case for Linked Data in archives, libaries and museums.

Friday, 23 March 2012

UKAD Conference

I gave a talk with Robert Baxter from Cumbria Archive Service at the annual UKAD conference at The National Archives on 21 March. The talk explored the potential of Linked Data in the archives, libraries and museums sector, focusing on the experience of Step change.

Key lessons/challenges from Step change and from other JISC projects King's College Archives are working on (Trenches to Triples and World War One Research) mentioned in the talk include:

  • Definining/setting up/maintaining APIs - this is potentially challenging and time-consuming
  • Need for URI definitions/syntax across the archives, libraries and museums sector - this discussion was started by LOCAH and is ongoing. A wiki will be launched soon by ULCC inviting information professional feedback on these definitions and to try and reach some consensus in the coming months
  • Place name vocabularies are a particular challenge. Step change archivists will potentially have access to some four or five sets of data about similar places - for example Geonames, English Place Names, AIM25-UKAT, GoGeo, and a local CALM place dataset. Have will they ensure consistency or that terms found across the datasets are actually talking about the same place?
  • Linked Data analysis exposes poor quality and inconsistent existing metadata. Step change is partly about providing tools that will identify discrepancies and make metadata input more consistent but the funding and management challenges of this laundry operation remain considerable
  • Establishing and supporting new live LOD services beyond existing JISC funding will be a challenge. Services go down - how will data retrieval cope with this fact of life?
  • Visualisation - this poses multiple challenges. how much information is too much information for users? How do we maintain relevancy - can the users decide themselves to some extent?
The development of the Workflow tool (Alicat) and LOD version of UKAT are well under way (Workpackages 2-3). These are informing redesign currently under way at Axiell. A meeting is planned on 29 March to review CALM development to date, prior to the commencement of analysis of Cumbria test catalogues by Robert using the new tools in a CALM development environment.