Friday, July 31, 2009

Introduction to Metadata - Online

Introduction to Metadata (online book)

Stephen Downes has an entry about a free online book about metadata - and provides a short review that finds the first three chapters comprehensive, but the treatment of "rights metadata" poor.

Nonetheless, worth looking at.

Introduction to Metadata. Version 3.0, edited by Murtha Baca. "An online publication devoted to metadata, its types and uses, and how it can improve access to digital resources." Published by Getty.edu.

Interesting note: "Reader's Note: The editor and authors of this publication are aware that the noun "metadata" (like the noun "data") is plural and, therefore, should take a plural verb form. However, in order to avoid awkward locutions, it has been treated here throughout as singular."

Contents:

Home
Introduction
Setting the Stage
Metadata and the Web
Crosswalks, Metadata Harvesting, Federated Searching, Metasearching
Rights Metadata Made Simple
Practical Principles for Metadata Creation and Maintenance
Glossary
Selected Bibliography
Contributors

Web 3.0 a matter of restructuring

Web 3.0: The Next Step for the Internet by Michael Baumann, Information Today (May 2009) via Allbusiness.com

Web 3.0, says Nova Spivack, CEO of Radar Networks, is going to be about restructuring the web.

"Web 3.0 will usher in a revolution in the construction of the internet itself. In the third decade of the Web, the focus is going back to the back end and we're focusing on upgrading the infrastructure of the Web again."

"When an application sees a page on the Web today it doesn't really know what to do with it," Spivack says. "But as we add more of this open standard metadata to the Web, it makes the Web machine understandable. Also, as applications get smarter because they can understand language and know what words mean, that also adds meaning and structure to the Web."

Wednesday, July 29, 2009

Resources for Thesaurus Construction

A posting in Buslib-L titled 'Summary of responses re thesaurus for records management purposes' (July 22) gives us a good starter list of resources for constructing thesauri. Thanks to Sarah Knight



Thesaurus Construction and Publishing Solutions

http://www.multites.com/index.htm
Provides software. Also has a list of Resources for Thesaurus Construction.


Taxonomy Warehouse (Factiva)
List of all vocabularies: http://www.taxonomywarehouse.com/querybyvoc_search_include.asp

"Taxonomy Warehouse was created in 2001 as a valuable community resource, available free to users and vocabulary publishers to help organizations maximize their information assets and break through today’s information overload."


Index New Zealand Thesaurus
http://innz.natlib.govt.nz/content/thesaurus/index.htm

"The Index New Zealand Thesaurus has been designed to describe journal and newspaper articles about New Zealand and the South Pacific in the areas of social sciences and humanities. This version was released in November 2005 and contains over 1,200 preferred terms."



Australian Governments' Interactive Functions Thesaurus (AGIFT)

http://www.naa.gov.au/records-management/create-capture-describe/describe/agift/agift-zip.aspx

"The Australian Governments’ Interactive Functions Thesaurus (AGIFT) is a three-level hierarchical thesaurus that describes the business functions carried out across Commonwealth, state and local governments in Australia. It contains 25 high-level functions, each with second and third level terms, as well as non-preferred terms and related terms. A scope note describes the range of activities covered by a preferred term and provides cross-references."


Victoria Online Thesaurus (July 2008)
Department of Innovation, Industry, and Regional Development
http://www.egov.vic.gov.au/index.php?env=-innews/detail:m2110-1-1-8-s-0:n-9-1-0--

"The Victoria Online Thesaurus is a subject thesaurus of descriptive terms that reflect the themes and resources within Victoria Online (VO), the Victorian Government’s online portal. Victoria Online is a metadata-driven gateway to Victorian State, Federal and Local government information. The VO Thesaurus has been developed to populate the Keyword (DC.Subject) field within the VO Metadata Application Profile (VOMAP)."


Health Thesaurus - Health and Ageing Thesaurus
Australia - Department of Health and Ageing
http://www.health.gov.au/internet/main/Publishing.nsf/Content/health-thesaurus.htm

"The Health and Ageing Thesaurus is a living working tool which assists consistency and subject retrieval of health and ageing concepts. By standardising concepts to one single subject heading, the Thesaurus forms the basis for a common terminology within the Department.

MeSH (medical Subject Headings) produced by the US National Library of Medicine has been used as the basis of the medical terms and the corresponding hierarchical schedules in the Health and Ageing Thesaurus. We are very grateful for their permission to use MeSH in this way. For this edition 2004 MeSH has been used."


National Taxonomy of Exempt Entities
http://nccs.urban.org/classification/NTEE.cfm

The NTEE-CC classification system is used by the IRS and NCCS to classify US nonprofit organizations. It divides the universe of nonprofit organizations into 26 major groups under 10 broad categories.


Google Book Search: Can use Google Book Search to find thesaurus - combine that term with others related to organizations, fund raising, charity etc.


TIPS: Taxonomies in the Public Sector
Taxonomies and thesauri: a list of references and resources for public
sector applications (Great Britain)
http://www.govtalk.gov.uk/documents/Bibliography2005-05-11.pdf

"This bibliography has been prepared for the Taxonomies in the Public Sector (TIPS), a discussion group which supports the Metadata Working Group by encouraging information professionals in the public sector to meet and develop guidance on the implementation of taxonomies and metadata. While IPSV (Integrated Public Sector Vocabulary) and its predecessors GCL and LGCL are the main focus, public sector applications commonly use a complementary specialised vocabulary in tandem with IPSV. the bibliography gives background references across the gamut from development to exploitation and sharing the outputs"


United Nations Bibliographical Information System Thesaurus
http://unhq-appspub-01.un.org/LIB/DHLUNBISThesaurus.nsf

Entries are in six languages.


UNESCO Thesaurus
http://www2.ulcc.ac.uk/unesco/

"The UNESCO Thesaurus is a controlled vocabulary developed by the United Nations Educational, Scientific and Cultural Organisation which includes subject terms for the following areas of knowledge: education, science, culture, social and human sciences, information and communication, and politics, law and economics. It also includes the names of countries and groupings of countries: political, economic, geographic, ethnic and religious, and linguistic groupings."


International Labour Organisation (ILO) Thesaurus
http://www.ilo.org//thesaurus/defaulten.asp

Can use the rotated index to get a sense of the categories and content.

Tuesday, July 28, 2009

Structure First

The Perfect Search By Penny Crosman, Intelligent Enterprise (March 1, 2006 )

Overview article on the importance of structuring content to improve findability of information in an enterprise.

"Google-style search is all right for some, but greater accuracy in the enterprise demands a mix of techniques including content tagging and taxonomy development and technologies such as entity, concept and sentiment extraction tools."

Sunday, July 12, 2009

Text Mining for Meaning

Many are pointing to the emergence of a Web 3.0 that brings together Web 2.0 collaborative qualities with a semantic web of connections and semantic tools for understanding.

Nstein, the digital publishing technologies company, held a webinar titled From Metadata to Meaning: Intelligence in the Semantic Era on June 25, 2009, moderated by Diane Burley.

Seth Grimes of Alta Plana presented the case of an exploding digital universe where individuals are creating 70% of the content. Search based on keyword matching won't be sufficient. The next generation of tools must make sense of the data to deliver meaningful answers. Text mining combined with semantic analysis is one direction. Google has some new capability at picking out data and its context - eg the employment data for a US state. Newssift from the Financial Times shows more capability at identifying organizations, places, people and topical themes in its aggregation of news. Grimes noted that semantically enriched content, linked data, context sensitivity, and location awareness are now recurring themes. Text mining / analytics is enabling Web 3.0 and the Semantic Web.

For designers and users this means:

• Automated content categorization and classification.
• Text augmentation: metadata generation, content tagging.
• Information extraction to databases.
• Exploratory analysis and visualization.

Jean-Michel Texier, Chief Technology Officer at Nstein Technologies, described Nstein modules for mining and analyzing data - essentially to make sense of vast stores of text by extracting concepts, identifying and normalizing entities, categorizing content, extracting relevant sentences, analyzing sentiment, and finding similar content.

The presentation is available in PDF on the dowload section of the Nstein website: http://www.nstein.com/en/downloads.php?sourceId=236 - registration required.

Friday, July 10, 2009

Semantic Metadata

Semantic Metadata & Sagacious Serendipty by Diane Burley, Silicon Valet (June 2009)

Digital media consultant, Diane Burley, says that metadata is the key to creating websites that enable discovery and provide user satisfaction. But she doesn't mean that run-of-the-mill metadata - date, author, format etc - administrative metadata, but semantic metadata.

"The academics at Kent State call it descriptive metadata, while the folks at Nstein prefer to call it semantic metadata (semdata??). It is metadata that is generated using a multi-faceted approach of computational and linguistic analysis. It not only extracts meaning from documents – but also embeds the synonyms, summary, categories, even the tone, in order to create a linguistic fingerprint. This linguistic fingerprint can then be matched against any other linguistic fingerprint – to find like pieces of content."