Tuesday, February 09, 2010

Taxonomy Boot Camp - Call for Speakers

Taxonomy Boot Camp will be November 15 to 16, 2010 - Making it Real: Getting Value, Support & Usage from Taxonomies. The organizers have issued a call for speakers to submit proposals for sessions. Deadline for submission is April 24.

Sunday, January 17, 2010

A hybrid semantic Web MDM approach

The Semantic Web: A Perfect Complement to Master Data Management Initiatives by J. Brooke Aker, Information Management (Jan 11)

To manage information this article proposes a hybrid of the "master data management" for defining business entities and semantic technologies to uncover unstructured data.

"Implementation-wise, semantic Web technologies allow IT managers to integrate master data without needing to understand the data model or writing complex SQL statements. Additionally, the semantic Web offers quick and precise data analysis for enterprise users to manage and discover relationships among master data based on semantic modeling and reasoning."

Friday, December 11, 2009

Designing for Faceted Search

Special Report – Designing for Faceted Search By Stephanie Lemieux (with Seth Earley & Associates) Altsearchengines (Dec 2009)

This article, originally published in KMWorld March 2009, is reproduced with diagrams and illustrations in AltSearchEngines. Article has good, practical advice for designing for faceted search.

Wednesday, December 09, 2009

Portfolio of articles on Taxonomies and Tagging

The FUMSI Folio on Taxonomies and Tagging looks to be a useful reference for taxonomists. This is a collection of articles on taxonomies and tagging practices published in October 2009. Only $64 US.

Table of Contents:

* Editorial Introduction by Karen Loasby, Contributing Editor, Manage
* Taxonomies and Tagging Survey Results
* Creating User Centred Taxonomies: Part One, by James Kelway
* Creating User Centred Taxonomies: Part Two, by James Kelway
* Folksonomies: Business Use, by Fran Alexander
* Automatic Classification: A Panel Discussion, by Karen Loasby
* Image Findability: Improving through Tags, by Ian Davis
* Becoming a Taxonomist: Real Life Stories, by Karen Loasby
* Recommended Resources

Friday, November 27, 2009

Digital Asset Management and Metadata for Images and Video

This piece was previously posted on - 'Ian Davis - managing information'

Missing out on the recent Photo Metadata Conference - http://bit.ly/6PlLJj - has reminded me how much I love working in the DAM world, in particular in the area of creating metadata and controlled vocabularies to support digital image and video search and browse.

Reading about the Photo Metadata Conference programme
it seems like there were some great presentations. I downloaded them all, they're available from the conference website, and had great fun going through all the excellent experiences, comments and ideas.

I wish I'd been there for Madi Solomon's keynote on the collapse of boundaries in the digital world. I agree that it's less and less about what format an asset is in and more about what that asset is, and how it needs to be organised to support its use.

Assets need to work for their places in the world. Finding them and using them needs to be simpler, and metadata and controlled vocabularies need to support and enable this.

Understanding the assets an organization has, analysing the needs of that organisation, and ensuring they have what they need and that each asset is organised to support its use, is where the really exciting and satisfying work is for me.

After having worked for Corbis from 1991 to 1999, in the early research and development days of digital image organisation and sale, I was excited to see Max Wieberneits presentation on still and video metadata.

Video and still images have much in common. I've blogged about this in the past and it's still a big area for me. Both asset types have technical metadata, depicted content metadata and aboutness metadata, to name but a few. Add to this the sound tracks for video - which can be indexed for retrieval, and the ability to segment video into scenes and key frames, and you have an exciting mix of metadata across both formats.

I agree with Max that using established metadata systems makes a huge amount of sense, as does working to get as much metadata as possible from the creators or custodians of images and video - it's much easier to capture metadata early on in the creation process than down the line, and some metadata will be lost if you leave its capture too late.

As Max says, one key concern for image and video asset metadata is the users of the assets. Different people have different needs and need different metadata. For many people a good level of access to video can be built using initial metadata associated with the videos, key scene and frame analysis and the indexing of the audio tracks of the videos. Whereas for others, access to the mood of the video may only come through music analysis, lack of noise at key moments, and manually applied subject tags.

On the image side, as Max says, editorial users have somewhat differing needs to commercial users of stock photos. Max showed a great slide listing a long set of conceptual keywords: 'comfortable, dreaming, luxury, spoiled' etc. I remember the fun we had creating these concepts, arranging them in hierarchies, providing synonyms for them, and creating definitions and application rules to control how they're assigned. It sounds easy, but trying to accurately use a concept like, "spoiled" or "luxury" often brings many challenges.

I've already touched on the needs of video users, and some of the basic ways video can be organised. It was great to read Lionel Faucher's piece on how a video agency uses metadata. Video is easier than still images to work with, automated solutions are more applicable to video and much more successful, but challenges still abound, as Lionel clearly shows in his presentation.

One of the interesting topics I've been following for a while is the metadata being generated from digital cameras, and the work being done to make more use of it. Related to this is the exciting area of geographic coordinate metadata, which is created by some digital cameras when a photo is taken, and the uses to which that can be put.

Two presentations in the area of geography and image metadata were given by Bern Beuermann
, and Ross Purves. A great research area was mentioned by Bernd - the taking of GPS co-ordinates and linking them to points of interest that are within a certain range of a GPS location. This can make the tagging of images with key depicted buildings, or topography a little easier and will produce many advantages for image tagging and retrieval..

A couple of things that I'm interested in were missing from the conference. I'd have liked to have seen more on: working with video soundtracks, automatic scene and frame analysis, and the place of manually applied tags in video indexing. I'd also like to have seen more about the creation of hybrid image retrieval systems that bring together content based image retrieval with controlled vocabulary and folksonomy tags. Maybe that's all for next year!

There also seemed to have been a big emphasis on technology, file formats, and metadata standards - in many ways the building blocks or key tools for organising and providing access to video and image content. What I'd have liked to see more of is the uses to which these building blocks have been put, the real world sharing of user needs and the challenges of actually making the technology and the supporting structures work to achieve business aims.

I should end by thanking the organisers of the event, and the presenters, for putting so many presentations online - it's very helpful and refreshing to have such a good level of access to this form of content.

One way in which I keep involved in the image and video world is through my involvement in the DAM Foundation on Linkedin. There is a coffee meet-up organised for this afternoon, which I hope will kick start a lot of exciting developments. I'll post more about the outcome of the meeting next week.

Ian

Monday, November 23, 2009

Taxonomies and new technology

The Death of Taxonomies, revisited by Theresa Regli, CMS Watch (Nov 13)

Technology of text mining, entity extraction and semantic analysis is doing the grunt work of taxonomists. Theresa Regli foresees change in the life and work (and even name) of the taxonomist.

"Taking taxonomies beyond what technology can achieve on its own is the metadata architect’s challenge for the next decade, because technology is at the point where it achieves what taxonomists were doing a decade ago."

Sunday, November 22, 2009

Nstein Semantic Search

Nstein Technologies Launches Semantic Site Search, press release, NStein (Nov 17)

"Nstein Technologies Inc. www.nstein.com (TSX-V: EIN), a leader in digital content management solutions for information-rich enterprises, today announced the release of a new product, Semantic Site Search (3S). “3S is a front-end, multi-index search engine designed to provide users an unparalleled search experience,” said Nstein CTO Jean-Michel Texier. 3S leverages Nstein’s patented text-mining technology to power a faceted site search which returns highly accurate results that are organized categorically."

Thursday, November 12, 2009

Text Analytics to Help in Classification

Rise of the Machines: The Role of Text Analytics in Record Classification and Disposition by James Santangelo, Information Management, ARMA (Nov/Dec 2009)

Classification is essential but may be overwhelming to staff. Because of the volume automated classification is needed - and text analytics software can help.

"The latest advancements in text analytics use sophisticated techniques to determine the conceptual meanings within each file to compensate for shortcomings and extend the functionality of the applications that use policy rule engines. Use of text analytics greatly increases the accuracy of the classification by interpreting the meaning of terms in their context instead of being limited by the character strings inherent in policy rule engines."

Text Analytics

Text Analytics Gains a Broader Audience in the Enterprise by Paula J. Hane, IT Newslinks (Nov 2)

Text analytics is becoming more important to search. As this article explains:

"Text analytics extracts key information from unstructured text and helps to retrieve otherwise hidden information. It is a key component of many customer relationship management (CRM) applications, as well as for media and publishing, competitive intelligence, reputation monitoring, e-discovery, compliance, and financial analysis. Because of this, we've seen a number of acquisitions of text analytics firms by larger search companies (Business Objects acquired Inxight, Reuters acquired ClearForest, SAS acquired Teragram, and IBM acquired SPSS) and an increased pace of product and service rollouts."

It's really automatic tagging. Susan Feldman said of one vendor, TEMIS:

""It's clear that text analytics has taken off as a hot market, and TEMIS' expansion of its US business underlines this fact." "As the volume and flow of information increases, publishers and corporations are turning to automation to tag their content to make it findable, to understand what their customers are saying, to monitor trends and opinions about their products and their companies. That's impossible, given the exponential growth of information that needs to be processed, unless the process is automated." "

Wednesday, November 11, 2009

OpenCalais is amazing

Learn About and Try OpenCalais (a Free Service from Thomson Reuters), ResourceShelf (Nov 6)

OpenCalais is making a dent in use of semantic technology to extract entities and topics from text.

"In a nutshell, OpenCalais uses semantic technology and natural language processing to analyze text and add metadata by drawing out entities from documents, blog posts, news stories, etc. In some cases, ths type of data can identify or help identify relationships between people, businesses, etc."

This post gives an example of what it can do, and points to the OpenCalais viewerbox where we can try it for ourselves - take a substantial story from an online news site and see the types of data that Calais can extract and organize.

Explore to see the power of the tool. Will we need taxonomies if we have tools like OpenCalais?

Further, we can have this at our fingertips for content we follow with Feedly, a Firefox plugin.

Feed(ly)ing The Enterprise
, by Jennifer Zaino, Semantic Web (Nov 9)

"For one thing, it’s the semantic technology embedded within Feedly, which uses the OpenCalais web service to get a clean representation of metadata behind content. That gives power to enterprise users such as marketing professionals, who might be subscribed to various blogs and feeds and services and different content that’s relevant to their brand."