Research Group

Linguistic Engineering Group

The Linguistic Engineering Group (Pol. Zespół Inżynierii Lingwistycznej; ZIL) works on multiple aspects of Natural Language Processing, with a particular focus on information extraction, semantic text processing, and corpus linguistics for Polish and beyond.

The Group develops widely used resources and tools — including the National Corpus of Polish (NKJP), the Polish Dependency Treebank, the Grammatical Dictionary of the Polish Language, terminology extractors TermoPL and TermoUD, and open-source systems such as LAMBO, COMBO, Morfeusz, and Korpusomat — that underpin tagging, parsing, and corpus analysis pipelines for Polish.

ZIL participates in the CLARIN-PL and DARIAH-PL research infrastructures and COST actions (currently as Grant Holder of UniDive), and runs numerous national and international projects funded by sources including NCN, NCBR, FNP, NAWA, NPRH, CEF, Horizon 2020, and DIGITAL.

Portrait of Łukasz Kobyliński

Group Head

Łukasz Kobyliński, PhD

Łukasz Kobyliński the Head of the Linguistic Engineering Group at Institute of Computer Science, Polish Academy of Sciences, focusing on the area of machine learning in NLP, corpus linguistics and document understanding.

Faculty Members

Meet the faculty members contributing to research in the Linguistic Engineering Group.

Publications

Browse publications authored by faculty members affiliated with the Linguistic Engineering Group.

View all publications

Selected Projects

Explore current initiatives led by the Linguistic Engineering Group. Each project demonstrates how cutting-edge language technology is deployed to solve complex analytical challenges.

PLLuM
Ministry of Digital Affairs grant

PLLuM – Polish Large Language Model

The consortium’s goal is to develop the first open Polish LLM and an associated smart assistant. The project will adhere to ethical and responsible best practices in AI, incorporating data representativeness, transparency and fairness.

Lead: Maciej Ogrodniczuk