Mishkah
مشكاة — "niche of light"A Quran companion: full Uthmani-script text across all 114 surahs, a Hifz memorization tracker with streaks, a live sensor-driven Qibla compass, and Adhan-based prayer-time notifications.
AI-Native Android Engineer — I build production Android apps in Kotlin, then build the evaluation pipelines that keep the AI inside them honest. Nine shipped apps, one growing practice in LLM-as-a-Judge design.
I started as an Android engineer shipping production apps — Kotlin, MVVM/MVI, Retrofit, Room, WorkManager, Firebase. Somewhere along the way, evaluating whether an LLM's answer could be trusted became just as core to the job as making a RecyclerView scroll smoothly.
Today I split my time between building intelligence-first Android apps and designing LLM-as-a-Judge systems — custom rubrics, chain-of-thought and few-shot judges, G-Eval scoring, adversarial test sets — inside Langfuse. I care about the same thing in both halves of the job: does this actually work for the person using it, or does it just look like it does.
Outside of client work, I write about how AI tools actually perform day to day, and I keep a set of small Android repos where I re-derive patterns (Clean Architecture, Hilt, MVI) from scratch so I understand them past the tutorial.
Kotlin · Jetpack Compose & Views · MVVM / MVI · Retrofit · Room · WorkManager · Firebase Crashlytics
Langfuse tracing · LLM-as-a-Judge · Prompt engineering (CoT, G-Eval) · Hallucination detection · AI safety & QA
Claude & Cursor for AI-assisted engineering · Python · Git/GitHub · Docker Compose · Postman
Independent builds spanning faith tech, fintech risk, wellness, games, and computer vision — all native Android, all shipped past the prototype stage.
A Quran companion: full Uthmani-script text across all 114 surahs, a Hifz memorization tracker with streaks, a live sensor-driven Qibla compass, and Adhan-based prayer-time notifications.
Fintech · HackathonAn offline, AI-assisted risk-profiling prototype for digital-wallet KYC onboarding — a five-analyzer rule-based risk engine producing plain-language Low/Medium/High explanations, built for a Mobilink-sponsored hackathon track.
Game · Android360 puzzles across 10 themed packs and 3 difficulties, a Daily Doodle challenge, and a pass-and-play Duel mode — every asset is original Canvas-drawn art, and the soundtrack is synthesized at runtime.
Reads context — calendar gaps, phone usage, sleep, mood inferred from voice prosody — to surface 2–5 minute habit windows, with an on-device, offline empathetic AI coach.
Scans fridge and pantry contents with CameraX + ML Kit, suggests waste-reducing recipes across 8 dietary filters, and hands off to grocery delivery when you're out.
A personal "Journey" through life-stage transitions, a ranked content feed, vision boards, and a private app-lock screen for a quieter kind of self-care app.
Browse and search verified official brand storefronts and social links worldwide, with favorites and a "near you" restaurants & delivery view filtered by country.
An on-device styling assistant using age- and gender-aware vision models to suggest outfits from what's actually in your wardrobe.
An A–Z alphabet trainer paired with a spelling "garden" game across Animal Friends, Nature Walk, and Fun Things categories, with voice and haptic feedback for toy-like taps.
The other half of the job: not building the feature, but proving the AI feature can be trusted before it ships.
End-to-end tracing, prompt versioning, and model comparison across multiple production AI applications.
Custom scoring rubrics, chain-of-thought and few-shot judges, and G-Eval techniques — then tracking whether the judge itself is reliable.
Curating evaluation datasets with golden answers and deliberately adversarial, edge-case inputs to stress-test model behavior.
Applying hallucination-detection and AI-safety practices across evaluation workflows to raise production reliability.
Running statistical analysis over judge outputs to drive concrete prompt and model improvements, not just intuition.
AI LinkedIn Manager and Femverse — evaluation pipeline design and judge construction for production AI features.
Published in Stackademic on Medium — mostly practical comparisons and guides, written from actually using the tools.
Smaller, focused repos — each one isolates a single Android pattern until I trust it.
Clean Architecture + MVVM for API integration
Clean-Architecture-MVIUnidirectional data flow with MVI
Hilt-Dependency-InjectionHilt setup for DI and object graphs
Understanding-FlowKotlin Flow — emit & collect, from scratch
News_AppLatest headlines in a simple layout
Weather_AppReal-time, location-based weather
SnapShare-AppImage browsing & media permissions
Status_SaverSave WhatsApp status media locally
QR-ScannerLightweight QR scan & decode
Broadcast_ReceiversSystem events with live notifications
Camera_FeatureCapture photos via device camera
+ 7 more repos →Full profile on GitHub
Based in Rawalpindi, Pakistan. Open to Android engineering, AI-native mobile, and LLM evaluation work.