International Journal of
Information and Education Technology

Editor-In-Chief: Prof. Jon-Chao Hong
Frequency: Monthly
ISSN: 2010-3689 (Online)
E-mali: editor@ijiet.org
Publisher: IACSIT Press
 

OPEN ACCESS
3.9
CiteScore

IJIET 2012 Vol.2(4): 348-353
doi: 10.7763/IJIET.2012.V2.149

Tokenization as Preprocessing for Arabic Tagging System

Ahmed H. Aliwy

Abstract

Tokenization is very important in natural language processing. It can be seen as a preparation stage for all other natural language processing tasks. In this paper we propose a hybrid unsupervised method for Arabic tokenization system, considered as a stand-alone problem. After getting words from sentences by segmentation, we used the author’s analyzer to produce all possible tokenizations for each word. Then, written rules and statistical methods are applied to solve the ambiguities. The output is one tokenization for each word. The statistical method was trained using 29k words, manually tokenized (data available from http://testing.mimuw.edu.pl\~aliwy) from Al-Watan 2004 corpus (available from http://sites.google.com/site/mouradabbas9/corpora). The final accuracy was 98.83%.

Keywords

  • Arabic Tokenization
  • Arabic segmentation
  • Arabic tagging
149-T057

How to Cite

Copied

Ahmed H. Aliwy, "Tokenization as Preprocessing for Arabic Tagging System," International Journal of Information and Education Technology, vol. 2, no. 4, pp. 348-353, 2012. https://doi.org/10.7763/IJIET.2012.V2.149

Copyright & License

Copyright © 2012 by the authors. This is an open access article distributed under the Creative Commons Attribution License which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited (CC BY 4.0).

Article Metrics in Dimensions

Menu