status: shipping · Rawalpindi, Pakistan

Afifa Ali

AI-Native Android Engineer — I build production Android apps in Kotlin, then build the evaluation pipelines that keep the AI inside them honest. Nine shipped apps, one growing practice in LLM-as-a-Judge design.

9Android apps shipped
18Public repos
10Essays on applied AI
2026AI evaluation, full-time
01 · about

Two disciplines, one habit of verifying things.

I started as an Android engineer shipping production apps — Kotlin, MVVM/MVI, Retrofit, Room, WorkManager, Firebase. Somewhere along the way, evaluating whether an LLM's answer could be trusted became just as core to the job as making a RecyclerView scroll smoothly.

Today I split my time between building intelligence-first Android apps and designing LLM-as-a-Judge systems — custom rubrics, chain-of-thought and few-shot judges, G-Eval scoring, adversarial test sets — inside Langfuse. I care about the same thing in both halves of the job: does this actually work for the person using it, or does it just look like it does.

Outside of client work, I write about how AI tools actually perform day to day, and I keep a set of small Android repos where I re-derive patterns (Clean Architecture, Hilt, MVI) from scratch so I understand them past the tutorial.

Android Engineering

Kotlin · Jetpack Compose & Views · MVVM / MVI · Retrofit · Room · WorkManager · Firebase Crashlytics

AI / LLM Evaluation

Langfuse tracing · LLM-as-a-Judge · Prompt engineering (CoT, G-Eval) · Hallucination detection · AI safety & QA

Tools

Claude & Cursor for AI-assisted engineering · Python · Git/GitHub · Docker Compose · Postman

02 · experience

Where the work happened.

Jan 2026 —
Present

Associate AI Engineer

9D Technologies · Rawalpindi, Pakistan
  • Built end-to-end LLM evaluation pipelines in Langfuse for multiple production AI applications, covering quality, safety, and performance.
  • Designed LLM-as-a-Judge systems using custom rubrics, chain-of-thought and few-shot prompting, and G-Eval techniques.
  • Curated evaluation datasets with golden answers, adversarial examples, and diverse edge cases to stress-test model behavior.
  • Managed Langfuse tracing, prompt versioning, and model comparisons; tracked judge reliability with statistical analysis to drive prompt and model improvements.
  • Applied AI safety and hallucination-detection practices across evaluation workflows to improve production reliability.
Jan 2025 —
Dec 2025

Android Engineer (Associate & Intern)

9D Technologies · Rawalpindi, Pakistan
  • Built and shipped scalable Android apps using Kotlin, MVVM, Retrofit, and Room, with WorkManager background processing, push notifications, and Firebase Crashlytics — the production mobile foundation now applied to shipping AI-powered app features.
Nov 2022 —
Feb 2023

Career-Prep Fellow

AMAL Academy — Stanford University–funded fellowship · Online, Pakistan
03 · projects

Nine apps, nine different problems.

Independent builds spanning faith tech, fintech risk, wellness, games, and computer vision — all native Android, all shipped past the prototype stage.

Mishkah app iconFaith · Android

Mishkah

مشكاة — "niche of light"

A Quran companion: full Uthmani-script text across all 114 surahs, a Hifz memorization tracker with streaks, a live sensor-driven Qibla compass, and Adhan-based prayer-time notifications.

KotlinMVIRoomMaterial 3
TrustLens app splash screenFintech · Hackathon

TrustLens

Kifayat Ki Rah

An offline, AI-assisted risk-profiling prototype for digital-wallet KYC onboarding — a five-analyzer rule-based risk engine producing plain-language Low/Medium/High explanations, built for a Mobilink-sponsored hackathon track.

KotlinMVP128 unit tests
Wordoodle gameplay screenshotGame · Android

Wordoodle

hand-drawn word search

360 puzzles across 10 themed packs and 3 difficulties, a Daily Doodle challenge, and a pass-and-play Duel mode — every asset is original Canvas-drawn art, and the soundtrack is synthesized at runtime.

Jetpack ComposeMVICanvas
MiWellness · AI

Miora

AI micro-habit journaling

Reads context — calendar gaps, phone usage, sleep, mood inferred from voice prosody — to surface 2–5 minute habit windows, with an on-device, offline empathetic AI coach.

ComposeHiltOn-device DSP
PpUtility · CV

PantryPilot

predictive pantry & meal planner

Scans fridge and pantry contents with CameraX + ML Kit, suggests waste-reducing recipes across 8 dietary filters, and hands off to grocery delivery when you're out.

CameraXML KitGeofencing
PtWellness

Petalore

women's self-care companion

A personal "Journey" through life-stage transitions, a ranked content feed, vision boards, and a private app-lock screen for a quieter kind of self-care app.

KotlinMVIWorkManager
VtDiscovery

Vitrina

"Trove" — brand discovery

Browse and search verified official brand storefronts and social links worldwide, with favorites and a "near you" restaurants & delivery view filtered by country.

KotlinViews/XML
WaAI · Vision

WardrobeAssistant

wardrobe-bot

An on-device styling assistant using age- and gender-aware vision models to suggest outfits from what's actually in your wardrobe.

TFLiteOn-device ML
WgKids · Education

Word Garden

Learning Alphabets

An A–Z alphabet trainer paired with a spelling "garden" game across Animal Friends, Nature Walk, and Fun Things categories, with voice and haptic feedback for toy-like taps.

KotlinViews/XML
04 · evaluation practice

Making sure the model's answer is actually right.

The other half of the job: not building the feature, but proving the AI feature can be trusted before it ships.

tracing

Langfuse pipelines

End-to-end tracing, prompt versioning, and model comparison across multiple production AI applications.

judging

LLM-as-a-Judge design

Custom scoring rubrics, chain-of-thought and few-shot judges, and G-Eval techniques — then tracking whether the judge itself is reliable.

datasets

Golden & adversarial sets

Curating evaluation datasets with golden answers and deliberately adversarial, edge-case inputs to stress-test model behavior.

safety

Hallucination & safety QA

Applying hallucination-detection and AI-safety practices across evaluation workflows to raise production reliability.

analysis

Statistical reliability tracking

Running statistical analysis over judge outputs to drive concrete prompt and model improvements, not just intuition.

applied

Evaluated projects

AI LinkedIn Manager and Femverse — evaluation pipeline design and judge construction for production AI features.

07 · skills

The toolkit.

Android Engineering

KotlinMVVMJetpack ComposeRetrofitRoom DatabaseWorkManagerFirebase CrashlyticsMaterial Design

AI / LLM Evaluation

LLM-as-a-Judge FrameworksPrompt Engineering (CoT, G-Eval)Evaluation Dataset DesignHallucination DetectionAI Safety & QALangfuse (Tracing, Versioning, Experiments)

Dev Tools & Languages

PythonGit / GitHubClaudeCursorDocker ComposePostman

Professional

Critical ThinkingProblem SolvingCommunicationTeamworkOwnership
08 · education

Background.

2020 — 2024

Bachelor of Computer Science

University of Kotli, Azad Jammu & Kashmir
2018 — 2020

Pre-Engineering

Aga Khan Higher Secondary School, Hunza

Languages

EnglishProfessional
UrduNative
ShinaNative
09 · contact

Let's build something worth trusting.

Based in Rawalpindi, Pakistan. Open to Android engineering, AI-native mobile, and LLM evaluation work.