자연어처리 시스템

 

자연어의 이론적인 연구만이 아니라, 이것을 응용하여 실제로 사용하는 자연어처리 시스템을 만들려는 시도도 활발하다. 현재까지 개발된 자연어 시스템은,

등으로 분류된다.

자연어처리 시스템의 초기 성공의 예로서는 Winograd 의 SHRDLU가 유명하다. 이것은 화면상의 나무쌓기를 자연어인 영어로 조작한다. 그 후 정신분석을 행하는 ELIZA, 여행계획의 질의응답 시스템인 GUS, 영문해석 시스템인 MARGIE, 대화해석 시스템인 TOPLE 등 여러 가지 자연어처리 시스템이 개발되었다.

자연어로 데이터베이스를 조작하는 시스템의 개발도 활발히 진행되고 있다. SRI가 개발한 대규모 분산데이터베이스 조회 시스템인 LADDER, 관계모델(relational model)의 창시자인 코드(Codd)가 개발한 RENDEZVOUS, 지질학에 관한 질의응답 시스템 LUNAR, IBM이 개발한 질의응답 시스템 REQUEST, 항공편에 관한 대규모 관계형(relational) 데이터베이스를 취급하는 질의 응답 시스템 PLANE 등이 대표적이다.

1970년 이후에는 자연어처리 시스템이 상품화되기 시작했다. INTELLECT는 자연어를 취급할 수 있는 최초의 상품으로, 아티피셜 인텔리전스사(Artificial Intelligence Inc.)가 발매한 상품의 데이터베이스용 자연어 인터페이스이다. 프리 어소시에이트사(Free Associate Inc.)의 THEMIS도 자연어 데이터베이스 인터페이스로, 새로운 단어를 시스템에 추가할 수도 있다.  SAVVY는 PC용 자연어 인터페이스, STRAIGHT TALK는 워드프로세서용 자연어 인터페이스이다. NATURAL LINK는 메뉴 중에서 구(句)를 선택하여 복잡한 질문을 만드는 시스템이다. 다른 시스템처럼 질문을 하지 않는다는 것은 아니다. 최초에 Q&A라는 종합 소프트웨어 패키지는 워드프로세서, 데이터베이스에 대한 명령어를 자연어인 영어로 할 수 있도록 구성하여 성공한 시스템이다.

한국에서는 한국전자통신연구소(ETRI; Electronic & Telecommunication Research Institute)가 주도하는 정부 국책연구개발과제인, 지능형 컴퓨터에서 자연어 인터페이스에 관한 연구를 한국과학기술원 인공지능센터와 공동으로 한 바 있다.
 실용화된 시스템으로서는, 한국과학기술원 전산학과 최기선 교수팀에서 개발한 한국어 텍스트의 자동색인검출기 KAIS(Korean Automatic Indexing System)와, 자동철자검사 및 교정시스템(Hspe 11)이 자연어처리 시스템의 실용화로서는 효시이다.
 KAIS는 1991년에, Hspe 11는 1989년에 시제품이 완성되었다. 이 외에도 금성소프트웨어의 철자검사 시스템이 상용화되는 등, 1990년 하반기부터 이에 대한 연구가 번성하기 시작하였다.

자연어(영문) 처리 시스템

분   야

시스템명

개          발

기             능

질의 응답
시스템

SHRDLU

MIT

나무쌓기의 QA시스템

GUS

제롯스사

여행계획의 QA시스템

ELIZA

ditto

정신분석의 QA시스템

SCHOLOR

MIT

지리학습용 CAI

문제해결
시스템

STUDENT

MIT

산수 문제해결 시스템

Newton

MIT

물리 문제해결 시스템

Isaoc

텍사스대

지리 문제해결 시스템

문장해석
시스템

MARGIE

스탠퍼드대

영문해석 시스템

TOPLE

MIT

대화이해 시스템

LINGOL

MIT

영문해석 시스템

데이터베이스

검색

시스템

LADDER

SRI

대규모 분산 데이터베이스 조회 시스템

RENDEZOUS

CODD

관계 데이터베이스 조작 시스템

LUNAR

BBN

지질학에 관한 질의응답 시스템

REQUEST

IBM

데이터베이스 검색 시스템

PLANS

일리노이대

항공편에 관한 질의응답 시스템

상용

시스템

INTELLECT

AI

main flame의 데이터베이스용 자연어 인터페이스

SAVVY

에쿠스카리바
테크놀로지

개인컴퓨터용 자연어 인터페이스

STRAIGHT TALK

딕타폰

워드프로세서와의 자연어 인터페이스

THEMIS

프리 어소시에이트

VAX-Orode용 자연어 데이터베이스 인터페이스

SUPER-NATURAL

마이크로데이터

정보처리시스템과의 자연어 인터페이스

NATURAL LINK

AI

메뉴에 의한 자연어 DBMS 인터페이스

PEARL

코그니티브
시스템

지식베이스와의 자연어 인터페이스

 

 

자연어시스템의 구성

 

역사적인 Natural Language Understanding System

General Syntactic Processor (GSP)   SAD-SAM   BASEBALL   SIR   

STUDENT   ELIZA   LUNAR      MARGIE   SAM   PAM   LIFER

General Syntactic Processor (GSP)

A versatile system for the parsing and generation of strings in NL. GSP can directly emulate several other syntactic processors, including Woods' ATN grammar. It is not in itself an approach to language processing, but a system in which various approaches can be described. The basic idea is a chart that represents both the grammar and the input sentence as a modified tree (modified by making it binary and interchanging nodes and arcs). Chart manipulation mechanisms operate on the grammar itself.

SAD-SAM

[Lindsay, 1963] Syntactic Appraiser and Diagrammer -- Semantic Analyzing Machine. Programmed by Robert Lindsay in 1963 at CMU. It used an basic English vocabulary (1,700 words) and followed a context-free grammar. It parsed input from left to right, built derivation trees, and passed them to SAM, which extracted the semantically relevant information to build family trees and find answers to questions.

BASEBALL

[Bert Green, 1963] An information retrieval program with a large database of facts about all American League games over a given year. It accepted input questions from the user, limited to one clause with no logical connectives.

SIR

[Bertram Raphael, 1968]

Semantic Information Retrieval system, it was a prototype "understanding" machine, since it could accumulate facts and then make deductions about them in order to answer questions.

STUDENT

[Daniel Bobrow, 1968] STUDENT was a pattern-matching natural language program written by Bobrow as his doctoral thesis work at MIT. It could solve high-school level algebra story problems.

LUNAR

[William Woods, 1973] LUNAR answered questions about the rock samples brought back from the moon using two databases -- the chemical analyzes and the literature references. Specifically, it helped geologists access, compare, and evaluate chemical analysis data on moon rocks and soil composition obtained from the Apollo-11 mission. It operated by translating a question entered in English into an expression in a formal query language. The translation was done with an ATN parser coupled with a rule-driven semantic interpretation procedure.

MARGIE

[Schank, 1973] Meaning Analysis, Response Generation, and Inference on English system, developed at Stanford in 1973. Provided an intuitive model of the process of natural language understanding; see Conceptual Parsing above.

MARGIE consisted of three components. The first, the conceptual analyzer, converted English sentences into a CD representation using production-like rules called "requests." The middle phase was an inference system that accepted CD propositions and deduced facts from it, given the current system memory. Inference knowledge was represented in a semantic net. (So, for example, if told John hit Mary the system might infer Mary might get hurt.)

The final component was a text-generation module that took internal CD representations and converted them into English-like output.

SAM

[Schank] Script Applier Mechanism. Makes use of frame-like data structures called scripts, which represent stereotyped sequences of events, to understand simple stories. Prototype frames make it possible to use expectations about the usual properties of known concepts and about what typically happens in a variety of familiar situations to help understand sentences about those objects and situations.

PAM

[Wilensky, 1978] Plan Applier Mechanism. Understands stories by determining the goals that are to be achieved in the story and attempting to match the actions of the story with the methods that it knows will achieve the goals.

Plans are the means by which goals are accomplished, and understanding plan-based stories involves discerning the goals of the actor and the methods by which the actor chooses to fulfill those goals. In a plan-based story, the understander must discern the goals of the main actor and the actions that accomplish those goals.

PAM tries to determine the main goal and the D-goals that will satisfy the goal. It analyzes input conceptualizations for their potential realization of one of the plan boxes that are called by one of the determined D-goals. PAM utilizes two kinds of knowledge structures in understanding goals: named plans and themes (such as LOVE, which contain background information upon which predictions can be based that individuals will have certain goals)..

The distinction between plan-based and script-based stories is simple: in a script-based story, parts or all of the story correspond to one or more scripts available to the story understander; in a plan-based story, the understander must discern the goals of the main actor and the actions that accomplish those goals.

LIFER

[Hendrix, 1977] Built at SRI, it is an off-the-shelf system for building "natural language front-ends" for applications in any domain. It has a set of interactive functions for specifying a language and parser.