Thursday, September 20, 2007

Endeca Discovery Suite

Search gets even smarter, KMWorld (Sept 19, 2007)

Endeca is extending the capabilities of extend the Endeca Information Access Platform (IAP)with the Endeca Discovery Suite. Some of the improvements relate to tag extraction and tag-based visualization.

  • tag extraction capabilities to pull together and reveal common themes, concepts and entities from text-based reviews, blogs and posts for use in site navigation, search relevancy and search engine optimization;

  • tag-based visualization and navigation to complement static and dynamic site navigation, giving users more ways to explore and find desirable content and products

  • meta-relational capabilities to link different content types by common concepts, allowing people to dynamically summarize the user-generated content associated with any set of products

Tuesday, September 11, 2007

Freebase for Structured Content

Freebase.com might be the forerunner of the semantic web that has long been talked about. Ivor Tossell at the Globe and Mail described in A web that can read itself may be in our future (Sept 10)

"Freebase, like Wikipedia, is an open encyclopedia that most anyone can edit. But alongside each free-form article in Freebase, there are database fields for relevant hard data points. If the article is about a movie, you'll find fields for its release date, director, producer, screenwriters and so on. If the article is about a city, it will have fields for the city's population and location. If the article is about an artist, it will have a field for every one of that artist's works."

Freebase is an open project for building structured data applications using types and defined properties. Film will have one set or properties, geographic places another set. It will support complex queries.

From the FAQ - "Finally, while information in Freebase appears to be structured much like a conventional database, it’s actually built on a system that allows any user to contribute to the schemas—or frameworks—that hold the data. This wiki-like approach to structuring information lets many people organize the database without formal, centralized planning. And it lets subject experts who don’t have database expertise find one another, and then build and maintain the data in their domain of interest."

Understanding OWL

Web Ontology Language (OWL) and Semantic Web by Goutam Kumar Saha, Ubiquity Volume 8 Issue 35 (Sept 10, 2007)

In this ACM IT magazine article, Saha describes the web ontology language (OWL) which is a principal part of enabling a "semantic web". Has illustrations, examples and explanations of code.

"Web Ontology Language (OWL) is a language for defining and instantiating web ontologies (a W3C Recommendation). OWL ontology includes description of classes, properties and their instances. OWL is used to explicitly represent the meaning of terms in vocabularies and the relationships between those terms. Such representation of terms and their interrelationships is called ontology. OWL has facilities for
expressing meaning and semantics and the ability to represent machine interpretable content on the Web. OWL is designed for use by applications that need to process the content of information instead of just presenting information to humans. This is used for knowledge representation and also is useful to derive logical consequences from OWL formal semantics."

Card Sorting Challenges

Card Sorting: Mistakes Made and Lessons Learned By Sam Ng, UX Matters (Sept 10, 2007)

The author speaks from experience in this article about card sorting. It's a simple concept, deceptively so, and people may expect more than it can deliver.

"I’ve accepted the fact that card sort analysis—much like usability test analysis—is often messy and subjective. It’s part science, but mostly art. As with many aspects of our work, there isn’t necessarily a single correct, quantitative answer, but rather a number of different qualitative answers—all of which could be correct. Our job is to use our experience and our understanding of people to make judgment calls."

[Mentioned in InfoDesign: Understanding by Design ]

Wednesday, August 22, 2007

Debate about tagging

Is Tagging A Disruptive Innovation? - Joel Lamantia asks that question at Tagsonomy.com (July 21, 2007. The spread of tagging could distract from creating or maintaining taxonomies and possibly in use of metadata. But there could also be a large element of hype in the attention tagging is getting. This is one piece of a longer discussion. Lamnatia concludes "Though it’s been a few years since tagging became visible, it seems too early to understand what kind of changes - if any - will occur in the metadata management ecosystem as a result of tagging’s emergence."

One wonders if people really want to spend the extra few seconds to tag an item, and if they do tag to use something more useful than "read later". I suspect that tagging will remain personal, and that general access will depend on automatic categorization based on business rules.

Tuesday, August 21, 2007

Social search and taxonomies

What will be the impact on the use of taxonomies in companies as more adopt enterprise 2.0 ways of connecting people? Social search, social networking, social enterprise - these are new possibilities being adapted for enterprise use from the consumer universe of Facebook / MySpace, del.ico.us and other social bookmarking services, Digg and Flickr and all the places where people tag what they find. Ajay Gandhi at BEA Systems in a KMWorld webinar posits that social search tools will greatly enrich knowledge management by assisting in sharing knowledge and forming communities. Harvesting Enterprise Wisdom through Social Search reviews knowledge management archetypes, notes the rising state of information overload, and describes the ways social tools will help people cope with that load. It does not mean the end of formal taxonomies to support intranet navigation and repository search, but it will see the emergence of folksonomies based on how people tag and social networks developed according to interest and expertise.

Webinar will be available for 90 days at www.kmworld.com/webinars/bea/21aug2007

Thursday, August 16, 2007

Facets and Taxonomies

The Taxonomy Community of Practice (Earley and Associates) is running a session on Facets and Taxonomies Search online on August 29, 1pm to 2pm EDT. Price $50 US.

From the announcement: "We'll start with an overview of facets and faceted search and then hear from Peter Bell, one of the founders of Endeca, a faceted search company, about new developments in the field that allow a combination of unstructured and structured tagging and classification. "

Thursday, July 19, 2007

SchemaLogic's Content Tagging

SchemaLogic, an information management company, has released Business Semantics Management software specifically designed for media and publishing enterprises. Associated Press is among the first to adopt it. The software allows users to tag content while the software manages the semantic connections.

SchemaLogic provides a solution for customers to implement a collaborative process that enables writers, photographers, and editors to participate in the development and enrichment of the underlying “content tags” that describe information in a dynamic, ever-changing environment – and they do not have to change the way they use their own terminology. Content tagging is an advanced method of identifying and labeling information assets including audio, video, news stories, and other web content using text descriptions. SchemaLogic’s software manages the definition and relationships between content tags so that each individual in each department can continue to work in a way that makes sense for them, while the semantic differences are resolved by the technology.


SchemaLogic Delivers First Business Semantics Management Solution for Media and Publishing Enterprises, Press Release (July 16)

Friday, July 13, 2007

Information Professionals in the Text Mine By Kathryn A. Lavengood and Pam Kiser, Online (May / June 2007)

Authors argue that text mining - for drawing relationships between disparate data from many sources - needs a "semantic infrastructure that focuses on information quality and decision support".

Key point (bolding added) : "The interpretation of text is just the first step in making the information usable. Another key part is then organizing the resulting “text pieces” into some form of usable network. This is addressed by building taxonomies and ontologies that can be navigated to explore specific topics of interest. Finally, the results must be output in a format that can be interpreted and lead to knowledge discovery."

It identifies three parts to a text-mining system: parsing the text into parts, tagging extracted information, and organizing the parts using taxonomies and ontologies.

Wednesday, July 11, 2007

Dow Jones Releases Synaptica 6.4 for Improved Business Semantic Management

In early June 2007 Dow Jones & Company introduced Synaptica 6.4 - its latest semantic Web-enabled knowledge organization system for the enterprise.

Synaptica 6.4 simplifies and standardizes vocabulary and metadata management in order to unlock valuable business intelligence.

“Computers can store, search and display enormous amounts of information, but until recently machines have not been able to understand the meaning of the content,” said Dave Clarke, global taxonomy director, Dow Jones. “Now, with the semantic Web being able to capture the meaning in a machine-readable way, users can discover latent information and make new connections between isolated content while benefiting from comprehensive and precise information recall.”

To read the full press release visit http://www.factiva.com/investigative/releases/20070605_synaptica.asp?node=menuElem1176

For more information about Synaptica 6.4, visit http://www.factiva.com/products/taxonomy/synaptica.asp?node=menuElem1511

To learn more about Dow Jones services, visit www.dowjones.com/clientsolutions