Don't Just Scratch the Surface: Enhancing Word Representations for Korean with Hanja

08/25/2019
by   Kang Min Yoo, et al.
0

We propose a simple approach to train better Korean word representations using additional linguistic annotation also known as Hanja. Using its association with the Chinese, we devise a method to transfer representations from the language by initializing Hanja embeddings with Chinese ones. We evaluate the intrinsic quality of representations built upon our approach through word analogy and similarity tests. In addition, we demonstrate their effectiveness on several downstream tasks including a novel Korean news headline generation.

READ FULL TEXT

Please sign up or login with your details

Forgot password? Click here to reset