DeepX Hackathon 2026

Arabic Aspect-Based Sentiment Analysis

A production-grade NLP system that dissects Arabic customer reviews — identifying what aspects are mentioned and how the customer feels about each one. Built on MARBERTv2 with semi-supervised learning.

82.8% F1 Score
9 Aspects
3,584 Training Samples
3 Training Stages
Try It Live
Live Demo

Analyze Arabic Reviews

Type or paste any Arabic customer review and watch the model break it down into aspects and sentiments in real-time.

📝 Input Review

Try an example:

🎯 Analysis Results
🔍

Enter an Arabic review and click Analyze to see the results.

Technical Deep Dive

Model Architecture

How our system processes Arabic reviews from raw text to structured sentiment predictions.

1

Arabic Preprocessing

Normalize Alef/Taa forms, strip diacritics, collapse repeated chars, remove non-Arabic noise. Filter out reviews with <50% Arabic characters.

12% training data filtered
2

MARBERTv2 Encoder

Pre-trained on 1B+ dialectal Arabic tweets. 12-layer transformer producing 768-dim [CLS] representations. Fine-tuned end-to-end.

768-dim embeddings
3

9 Classification Heads

Independent linear layers (768→4) per aspect. Softmax forces mutual exclusivity: each aspect gets exactly one of [absent, positive, negative, neutral].

9 × 4-class softmax
4

Weighted Loss + Pseudo-Labels

Inverse-frequency class weights handle imbalance. Semi-supervised pseudo-labeling adds 1,418 high-confidence samples from unlabeled data.

3,584 total training samples
Pipeline Overview
Arabic Review
"الأكل كان ممتاز بس الخدمة سيئة"
Clean & Tokenize
Normalize + MARBERT Tokenizer
MARBERTv2
12-layer Transformer
9 Aspect Heads
Softmax Classification
Structured Output
food: positive, service: negative
Performance

Model Metrics

Evaluation results on the hidden test set — measured using Micro F1-score.

🎯
82.8%
Micro F1-Score
📐
81.3%
Precision
📡
84.4%
Recall
🏆
24.85/30
Pillar 1 Score

Training Progression

Stage 1 — Supervised Training

1,731 Arabic-filtered samples. Early stopping at epoch 7.

Val F1: 78.19%

Stage 2 — Pseudo-Labeling

2,972 unlabeled samples → 1,418 high-confidence pseudo-labels (≥90% confidence).

47.7% yield rate

Stage 3 — Full Retrain

3,584 combined samples (train + val + pseudo). 10 epochs, loss: 1.0 → 0.04.

Test F1: 82.8%
Data Insights

Prediction Analysis

Distribution of aspects and sentiments across all 500 test predictions.

Aspect Distribution

Loading chart data...

Sentiment Distribution

Loading chart data...

Aspect × Sentiment Breakdown

Loading chart data...
Taxonomy

Aspect Categories

The 9 aspect categories our model detects in every review.

🍽️

Food

الطعام

Quality, taste, freshness, portion size

🤝

Service

الخدمة

Staff attitude, speed, professionalism

💰

Price

السعر

Value for money, pricing fairness

Cleanliness

النظافة

Hygiene, tidiness of the venue

🚚

Delivery

التوصيل

Speed, packaging, accuracy

🎶

Ambiance

الأجواء

Atmosphere, decor, noise level

📱

App Experience

التطبيق

App usability, bugs, interface

📋

General

عام

Overall impression, recommendation

None

لا يوجد

No specific aspect mentioned