The Geographic Names Information System (GNIS) is a database of name and location information about more than two million physical and cultural features, encompassing the United States; the associated states of the Marshall Islands, Micronesia, Palau, and Antarctica. It is a type of gazetteer. It was developed by the United States Geological Survey (USGS) in cooperation with the United States Board on Geographic Names (BGN) to promote the standardization of feature names. Data were collected in two phases. A third phase was considered, which would have handled name changes where local usages differed from maps, it was never begun. The database is part of a system that includes topographic map names and bibliographic references. The names of books and historic maps that confirm the feature or place name are cited. Variant names, alternatives to official federal names for a feature, are also recorded. Each feature receives a permanent, unique feature record identifier, sometimes called the GNIS identifier. The database never removes an entry, "except in cases of obvious duplication."
Original purposes The GNIS was originally designed for four major purposes: to eliminate duplication of effort at various other levels of government that were already compiling geographic data, to provide standardized datasets of geographic data for the government and others, to index all of the names found on official U.S. government federal and state maps, and to ensure uniform geographic names for the federal government.
Phase 1 Phase 1 lasted from 1978 to 1981, with a precursor pilot project run over the states of Kansas and Colorado in 1976, and produced 5 databases. It excluded several classes of feature because they were better documented in non-USGS maps, including airports, the broadcasting masts for radio and television stations, civil divisions, regional and historic names, individual buildings, roads, and triangulation depot names. The databases were initially available on paper (2 to 3 spiral-bound volumes per state), on microfiche, and on magnetic tape encoded (unless otherwise requested) in EBCDIC with 248-byte fixed-length records in 4960-byte blocks. The feature classes for association with each name included (for examples) "locale" (a "place at which there is or was human activity" not covered by a more specific feature class), "populated place" (a "place or area with clustered or scattered buildings"), "spring" (a spring), "lava" (a lava flow, kepula, or other such feature), and "well" (a well). Mountain features would fall into "ridge", "range", or "summit" classes. A feature class "tank" was sometimes used for lakes, which was problematic in several ways. This feature class was undocumented, and it was (in the words of a 1986 report from the Engineer Topographic Laboratories of the United States Army Corps of Engineers) "an unreasonable determination", with the likes of Cayuga Lake (in upstate New York) being labelled a "tank". The USACE report assumed that "tank" meant "reservoir", and observed that often the coordinates of "tanks" were outside of their boundaries and were "possibly at the point where a dam is thought to be".
National Geographic Names database The National Geographic Names database (NGNDB hereafter) was originally 57 computer files, one for each state and territory of the United States (except Alaska which got two) plus one for the District of Columbia. The second Alaska file was an earlier database, the Dictionary of Alaska Place Names that had been compiled by the USGS in 1967. A further two files were later added, covering the entire United States and that were abridged versions of the data in the other 57: one for the 50,000 most well known populated places and features, and one for most of the populated places. The files were compiled from all of the names to be found on USGS topographic maps, plus data from various state map sources. In phase 1, elevations were recorded in feet only, with no conversion to metric, and only if there was an actual elevation recorded for the map feature. They were of either the lowest or highest point of the feature, as appropriate. Interpolated elevations, calculated by interpolation between contour lines, were added in phase 2. Names were the official name, except where the name contained diacritic characters that the computer file encodings of the time could not handle (which were in phase 1 marked with an asterisk for update in a later phase). Generic designations were given after specific names, so (for examples) Mount Saint Helens was recorded as "Saint Helens, Mount", although cities named Mount Olive, not actually being mountains, would not take "Mount" to be a generic part and would retain their order "Mount Olive". The primary geographic coordinates of features which occupy an area, rather than being a single point feature, were the location of the feature's mouth, or of the approximate center of the area of the feature. Such approximate centers were "eye-balled" estimates by the people performing the digitization, subject to the constraint that centers of areal features were not placed within other features that are inside them. alluvial fans and river deltas counted as mouths for this purpose. For cities and other large populated places, the coordinates were taken to be those of a primary civic feature such as the city hall or town hall, main public library, main highway intersection, main post office, or central business district regardless of changes over time; these coordinates are called the "primary point". Secondary coordinates were only an aid to locating which topographic map(s) the feature extended across, and were "simply anywhere on the feature and on the topographic map with which it is associated". River sources were determined by the shortest drain, subject to the proximities of other features that were clearly related to the river by their names.
… excerpt ends here. Continue reading the full article.


