On the resolution of ambiguities in the extraction of syntactic categories through chunking

Freudenthal, D; Pine, JM; Gobet, F

Please use this identifier to cite or link to this item: http://bura.brunel.ac.uk/handle/2438/802

Full metadata record

DC Field	Value	Language
dc.contributor.author	Freudenthal, D	-
dc.contributor.author	Pine, JM	-
dc.contributor.author	Gobet, F	-
dc.coverage.spatial	25	en
dc.date.accessioned	2007-05-25T09:09:39Z	-
dc.date.available	2007-05-25T09:09:39Z	-
dc.date.issued	2005	-
dc.identifier.citation	Cognitive Systems Research, 6(1): 17-25, Mar 2005.	en
dc.identifier.uri	http://www.sciencedirect.com/science/article/pii/S1389041704000610	en
dc.identifier.uri	http://bura.brunel.ac.uk/handle/2438/802	-
dc.description.abstract	In recent years, several authors have investigated how co-occurrence statistics in natural language can act as a cue that children may use to extract syntactic categories for the language they are learning. While some authors have reported encouraging results, it is difficult to evaluate the quality of the syntactic categories derived. It is argued in this paper that traditional measures of accuracy are inherently flawed. A valid evaluation metric needs to consider the wellformedness of utterances generated through a production end. This paper attempts to evaluate the quality of the categories derived from co-occurrence statistics through the use of MOSAIC, a computational model of syntax acquisition that has already been used to simulate several phenomena in child language. It is shown that derived syntactic categories that may appear to be of high quality quickly give rise to errors that are not typical of child speech. A solution to this problem is suggested in the form of a chunking mechanism that serves to differentiate between alternative grammatical functions of identical word forms. Results are evaluated in terms of the error rates in utterances produced by the system as well as the quantitative fit to the phenomenon of subject omission.	en
dc.format.extent	169878 bytes	-
dc.format.mimetype	application/pdf	-
dc.language.iso	en	-
dc.publisher	Elsevier	en
dc.subject	Distributional learning	en
dc.subject	Co-occurrence statistics	en
dc.subject	Syntactic categories	en
dc.subject	MOSAIC	en
dc.subject	Chunking	en
dc.subject	Language acquisition	en
dc.subject	Cognitive modelling	en
dc.title	On the resolution of ambiguities in the extraction of syntactic categories through chunking	en
dc.type	Research Paper	en
Appears in Collections:	Psychology Department of Life Sciences Research Papers

Files in This Item:

File	Description	Size	Format
freudenthal-CSR-all.pdf		165.9 kB	Adobe PDF	View/Open

Show simple item record