Software, Models & Datasets

The SVARNA research toolkit: open platforms, code, models and corpora

Open platforms, code, models and corpora, openly licensed where noted.

Applications & Platforms

Greek NLP Swiss Knife

A unified portal hosting 10 containerized AI applications (~200k lines of code) for Modern Greek linguistics, classical studies, and digital humanities research, developed with cloud computing support from Microsoft. Role: Principal Architect & Main Developer.

MuVeS

AI research assistant platform for paper discovery, automated literature reviews, interactive chat, paper analysis. Role: One of three main developers.

MEDEA-NEUMOUSA

AI platform for classical philology: (1) Translation between 18 ancient languages, (2) Knowledge graph extraction (RDF/TTL), (3) Zeugma neuro-symbolic reasoning (LLM+Prolog), (4) Emotion knowledge graphs, (5) Semantic analysis. Role: Main developer.

Plot Analyzer

Bidirectional neuro-symbolic narrative analysis platform with 8-stage hybrid architecture: neural perception via LLMs, symbolic reasoning based on 5 narrative theories (Aristotelian Poetics, Russian Formalism, etc.), bidirectional feedback loop, 14-type conflict taxonomy, and chunking & reconciliation for long texts. Role: Main developer.

Simasia-Studio (TextCraft)

AI text editor and translator with RAG for domain-specific translation, grammar/style analysis, track changes output. Role: Main developer.

RAG-to-Coq Pipeline

Historical event extraction with 10 extraction modes, 5 LLM support, translation to Coq for formal verification. Role: Main developer.

NATS

NLP analysis suite: document embeddings, NER (19 types), network analysis. Role: Main developer.

Linguistic Distance Calculator

7-dimension language distance measurement. Role: Main developer.

Greek Curriculum Ontology Extractor

LLM-based ontology extraction with RAG. Role: Main developer.

Syntax-Expert

Multi-framework syntactic analysis (Minimalism, HPSG, LFG, DS). Role: Main developer.

DI_detector

Greek dialect identification. Role: Main developer.

TextCraft Terminography

AI-assisted terminography platform: automatic term extraction from corpora (PDF/DOCX/CSV) per ISO 1087-1:2000, neuro-symbolic definition evaluation, RAG pipeline with OpenAI embeddings, multi-LLM support (Gemini, Claude, GPT-4o, DeepSeek), ELETO terminology scraper. Role: Main developer.

Svarna: Greek Corpus Workbench

Corpus linguistics workbench for Modern Greek: KWIC concordancing, frequency and n-gram analysis, discourse markers, keyness (log-likelihood), regex search, and LLM-assisted pragmatic analysis over 500M+ words across six corpora. Role: Main developer.

Greek Rhyme System

AI-powered identification and generation of rhyme patterns in Modern Greek poetry: full rhyme taxonomy (M/F2/F3, RICH, IDV, MOS, IMP), multi-LLM analysis with RAG corpus retrieval and a deterministic phonological verification loop. Role: Main developer.

Greek Dialect Generator

Text generation in four Greek dialects (Pontic, Cretan, Northern Greek, Cypriot) using LoRA adapters fine-tuned on GRDD+ (23k+ examples) over Llama 3.1, Llama 3, and Krikri base models, served with 4-bit quantization. Role: Main developer.

Voyant-NLP

Text analysis and embeddings lab: word frequencies, concordance, collocates, Word2Vec/FastText training with analogies, NER, POS tagging, topic modeling, sentiment analysis, and document clustering. Role: Main developer.

Code & Proof Assistants

Coq for NL Semantics / FraCoq

Proof assistant code for MTT semantics and NLI. Role: Main developer (MTT book), contributor (FraCoq).

Compositional Bayesian Semantics

Haskell implementations. Role: Contributor.

Anvec

Metaphoricity detection. Role: Contributor.

Fine-tuned Models

Krikri-8B Base LoRA

Fine-tuned Llama model for Greek dialectal varieties (Cretan, Cypriot, Northern, Pontic). Trained on GRDD dataset.

Llama-3 8B Instruct LoRA

Fine-tuned Llama-3 8B for Greek dialectal NLP.

Llama-3.1 8B Instruct LoRA

Fine-tuned Llama-3.1 8B for Greek dialectal NLP.

Datasets & Benchmarks

HeptaTax

Dataset and benchmark for classification of 16th-century Heptanesian notarial acts, with neuro-symbolic classification system and evaluation. Role: Principal contributor.

GRDD/GRDD+

Greek Regional Dialects: 11 varieties, ~7M words. Role: Main creator.

DNLI

First dialogue NLI with disfluencies. Role: Co-creator.

OYXOY

Greek NLU benchmark: NLI (1,763), WSD (6,896), metaphor (14,416). Role: Co-creator.

SuperOYXOY

Extended: paraphrase, augmented NLI, bias detection. Role: Co-creator.

Fine-Grained Entailment

Greek FraCaS extension (428), RTE annotation, Greek XNLI. Role: Main creator.

Precise Entailment

Expert-annotated NLI (150 examples). Role: Contributor.

Shami

Levantine Arabic: 110K sentences. Role: Co-creator.

ATSAD

Arabic Tweets Sentiment: 36K tweets. Role: Contributor.

Shami-Senti

Levantine sentiment (~2.5K). Role: Contributor.

Greek Rhyme Dataset

Dataset for Greek rhyme analysis. Role: Main creator.

Interwar Poetry & Prose

Modern Greek interwar poetry corpus for RAG generation. Role: Main creator.