Paper

Multilingual segmentation based on neural networks and pre-trained word embeddings

The DISPRT 2019 workshop has organized a shared task aiming to identify cross-formalism and multilingual discourse segments.
Elementary Discourse Units (EDUs) are quite similar across different theories. Segmentation is the very first stage on the way of rhetorical annotation. Still, each annotation project adopted several decisions with consequences not only on the annotation of the relational discourse structure but also at the segmentation stage.
In this shared task, we have employed pre-trained word embeddings, neural networks (BiLSTM+CRF) to perform the segmentation.

EusDisParser: improving an under-resourced discourse parser with cross-lingual data

Development of discourse parsers to annotate the relational discourse structure of a text is crucial for many downstream tasks.
However, most of the existing work focuses on English, assuming a quite large dataset.
Discourse data have been annotated for Basque, but training a system on these data is challenging since the corpus is very small.
In this paper, we create the first parser based on RST for Basque, and we investigate the use of data in another language to improve the performance of a Basque discourse parser.

Towards discourse annotation and sentiment analysis of the Basque Opinion Corpus

Discourse information is crucial for a better understanding of the text structure and it is also necessary to describe which part of an opinionated text is more relevant or to decide how a text span can change the polarity (strengthen or weaken) of other spans by means of coherence relations.
This work presents the first results on the annotation of the Basque Opinion Corpus using Rhetorical Structure Theory (RST).

Weighted finite-state transducers for normalization of historical texts

This paper presents a study about methods for normalization of historical texts. The aim of these methods
is learning relations between historical and contemporary word forms. We have compiled training and test
corpora for different languages and scenarios, and we have tried to read the results related to the features
of the corpora and languages. Our proposed method, based on weighted finite-state transducers, is com-
pared to previously published ones. Our method learns to map phonological changes using a noisy channel

Ipuin-moldaketa herri-hizkerara egokitzeko, aldatzeko eta modu esanguratsuan kontatzeko markaketa: Ahozko komunikazioa lantzen eta aztertzen Haur Hezkuntzako gelan

Lan honetan Martin Txiki eta Basajaunak ipuinaren irakurketa ozen esanguratsuak talde zehatz bateko haurren hizkuntza- eta komunikazio-gaitasunean izan duen eragin zuzena aztertu da, beren beregi diseinatutako esku-hartze eta ikerketa baten bidez. Zehatzago, ipuina bizkaierara moldatu da eta Mungiako Legarda HLHI ikastetxeko Haur Hezkuntzako 4 urteko gela batean irakurri da, horrek haurrek duten euskalkiaren ezagutzan eta egiten duten erabileran duen eragina aztertzeko. Esku-hartzea egiteko lan moldea konstruktibismoan oinarritu da.

Zer i(ra)kas dezakegu geure corpusekin "jolastuz"?

Hizkuntzak ikasteko askotariko metodologiak erabili izan dira: metodo zuzena, itzulpen metodoa, metodo audiolinguala, metodo komunikatiboa, hurbilpen lexikoa, ariketetan oinarritutako metodoa, ikasleen erroreetan oinarritutakoa edota metodo eklektikoak. Azken urteotan, berriz, corpusekin «jolasteak» hizkuntzak modu esanguratsuan i(ra)kasteko aukerak eskaintzen dizkigunaren ustean gaude.

Pages

Subscribe to RSS - Paper