It was not a single voice command. It was a real-time contextual layer.
MARI is an Android app (React Native + Kotlin) that listens to Spanish voice commands, interprets intent with local ML (ONNX/TFLite), and executes guided actions across WhatsApp, YouTube, and Maps with accessibility and notification context.
Real problem#
Traditional mobile assistants fail at useful contextual execution: they parse words, but not the active screen state, current notifications, or execution risk.
Solution#
I designed a contextual orchestration layer with smart confirmation and robust fallback (ML + rules), prioritizing precision and safety before executing actions in sensitive apps.
What the app is#
MARI is an Android contextual voice assistant app. It combines React Native product experience with Kotlin native capabilities to capture voice, classify intent locally, and execute safe actions based on runtime context.
How it works (real flow)#
Captures voice input via microphone (manual command or wake word).
Converts speech to text (native Android STT / Sherpa).
Classifies intent with local ML.
Extracts entities (contact, destination, message, etc.).
Applies decision policy (confirm, clarify, or execute).
Executes native action (open app, navigate, reply, search).
Speaks response (TTS) and updates contextual state.
Technical base#
UI and navigation: React Native + TypeScript with contextual screens.
Native layer: Kotlin for Android system capabilities.
Voice stack: TTS (react-native-tts), native STT + Sherpa, wake word (Sherpa/Porcupine).
Local AI: ONNX Runtime + TensorFlow Lite with rules fallback.
Local memory: Room DB for aliases, recent commands, corrections, and routines.
Local MLOps: Python scripts for dataset, training, and ONNX/TFLite export.
Modules and integrations#
UX layer: Home, voice-guided onboarding, live contextual panel.
Orchestration: intent decision, confirmations, clarifications, execution.
Usage context: Accessibility Service + Notification Listener.
App domains: WhatsApp, YouTube, calls, Google Maps.
Coordinated voice pipeline to avoid microphone conflicts.
Code references
App.tsx
HomeScreen.tsx
ContextAwareOrchestrator.ts
AssistantBridgeModule.kt
AndroidManifest.xml
Architecture#
Hybrid architecture: React Native for product iteration speed, Kotlin for system-level capabilities, on-device inference for low latency, and clarification policies for safer intent handling.
Strengths#
On-device inference with low latency and reduced cloud dependency.
Solid RN + Kotlin hybrid architecture for UX + native control.
Context-assisted mode aware of screen state and notifications.
Coordinated voice pipeline between wake word and manual command.
Robust fallback with clarification and confirmation policies.
Test coverage: 20 TypeScript tests and 23 Kotlin tests.
Result#
MARI improves control and accessibility in real mobile usage scenarios, with more reliable responses and lower operational friction for users who depend on voice and context.
Project summary#
I built MARI as a contextual Android voice assistant with a hybrid React Native + Kotlin architecture and on-device ONNX/TFLite inference. I delivered the full voice pipeline, app-aware contextual orchestration, accessibility/notification integration, Room-based local memory, and a local MLOps flow for training and packaging models.
If your team is hiring to build real AI-powered mobile products, I am interested in joining.
I am looking to contribute to serious projects where I can lead architecture and end-to-end execution.
