Papers by Pranavan Theivendiram
Improving domain-specific SMT for low-resourced languages using data from different domains (L18-1)
Copied to clipboard
| Challenge: | Evaluation of domain-specific statistical machine translation system for official government letters . use of pseudo in-domain data showed improvement for both test sets . |
| Approach: | They develop a statistical machine translation system for official government letters . the system is based on a parallel in-domain dataset containing official letters based in Sinhala and Tamil . |
| Outcome: | The proposed system improves on the in-domain data in the domain of official government letters . the evaluations show that the system requires quality data from diverse subject matters and sources to perform better. |