 |
Each month, nearly two million teachers, parents, and school administrators visit the Gateway to Educational Materials (GEM) on the Web to find learning objects (resources) quickly and effectively. [1] However, few visitors to GEM are aware of the research and development (R&D) that both define what GEM is today and chart where it will be tomorrow. In this article, I will talk about the current work of the GEM R&D team and show how that work will improve how GEM functions for teachers and learners around the world.
At the heart of what makes GEM work is what is called structured metadata. The term metadata has been defined quite simply as “data about data.” The notion of metadata is not new. When we search the bibliographic records in a card catalog or in an online public access catalog, we are searching metadata records. Each record or card consists of data describing some other data—like the data in a book or on a CD. Another example would be a manual or electronic address book. Each record in an address book consists of metadata (or statements) about a person or organization.
The difficulty with finding educational objects (or any particular object) on the Web has been the fact that no one has bothered to create concise structured metadata statements about those objects that would make them easier to find by search engines. One of the mandates of the GEM project is to facilitate the creation of such metadata by GEM Consortium members distributed around the world. That means designing tools to both catalog learning objects by generating metadata about them and to gather (or “harvest”) that metadata to build an easily searchable database. Much of the R&D work at GEM focuses on making the metadata creation and metadata searching tools more powerful, more efficient, and more effective. I will discuss several of these R&D projects in the following paragraphs.
Today, in the Web world in which we find ourselves, a useful system for retrieval of geographically distributed educational objects consists of certain fundamental components: (1) repositories of both the digital learning objects themselves and the metadata that describes those objects; (2) registries where information about the nature of metadata statements made about those objects are defined, maintained, and published; and (3) metadata tools necessary for the creation of the statements made about educational objects that will be stored in the metadata repositories and the metadata tools to search across those repositories for objects that meet the needs of the searcher.
Since the metadata record describing an educational object consists of a set of descriptive statements about that object (e.g., author, title, subject, and grade level), success in searching across those records is increased when the descriptive statements made are carefully and consistently constructed. One way to increase the level of search success is to use statement terminology drawn from carefully developed controlled vocabularies, thesauri, and taxonomies. For example, drawing subject terms to assign to a metadata record from well-known thesauri such as the Thesaurus of ERIC Descriptors, the Art and Architecture Thesaurus or the NASA Thesaurus or from a controlled vocabulary such as the Library of Congress Subject Headings increases the chances that objects that are topically similar will be consistently retrieved.
However, for both the person creating metadata records using terms from such controlled vocabularies and the end users wanting to search for useful educational objects across repositories of those metadata records, having digital access to the vocabularies, thesauri, and taxonomies is critical in order to assign terms to records and to select terms for purposes of searching. To date, there has been no consistent, standardized way for the creators and searchers of metadata to interact with digital repositories of these various forms of controlled vocabularies. One of the R&D goals of GEM is to develop a standard mechanism that will permit metadata creation and searching tools to interact with such repositories. The R&D effort is developing a set of standard communications protocols and schemas that define: (1) the content and sequence of the various messages traveling between the digital metadata creation and search tools and the geographically distributed digital repositories that contain the controlled vocabularies, and (2) the data structures that format the content of the messages for consistent interpretation by digital tools. Such standardized protocols and schemas will make it possible for a tool and a repository to communicate with each other even if they were unaware of each others' existence prior to the communicative act. As humans, we are able to communicate with total strangers because our speech acts consist of fairly well defined, although somewhat fuzzy, sets of socially acquired communications patterns (protocols) and data structures (languages and syntactic bindings). Unlike humans, machines cannot deal very effectively (if at all) with fuzziness; therefore, the protocols and data structures in machine communication must be very precisely defined, consistently structure |