doi: 10.7763/IJIET.2011.V1.63
Matching Entities by Their Thai and English Proper Names
Abstract
We have developed a framework to match records from different data sources that refer to the same entities, using proper names as the matching keys. There are two challenges to overcome. Firstly, there may be typographical errors. Secondly, some data sources may store the data in Thai characters while some store them in English characters. Thus, Thai-version keys are romanized and compared with English-version keys, using string comparators and rule-based decision function. We report our experimental results and problems encountered, as well as suggest future research directions.
Keywords
- Entity names
- record matching
- romanization
- Thai characters
How to Cite
Rangsipan Marukatat, "Matching Entities by Their Thai and English Proper Names," International Journal of Information and Education Technology, vol. 1, no. 5, pp. 384-388, 2011. https://doi.org/10.7763/IJIET.2011.V1.63
Copyright & License
Copyright © 2011 by the authors. This is an open access article distributed under the Creative Commons Attribution License which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited (CC BY 4.0).