Meet Koode Bot: The Offline Edge-AI Hospital Receptionist Revolutionizing Patient Triage

By Tech & Healthcare Innovation Desk
In an era where digital healthcare infrastructure is increasingly tethered to the cloud, reliance on high-speed internet and heavy server costs often creates deployment roadblocks in remote regions, underfunded clinics, and privacy-sensitive facilities. Enter Koode Bot, an ingenious, fully localized open-source hospital reception and triage system designed to function entirely offline.
By combining cutting-edge edge-AI hardware, localized Large Language Models (LLMs), automated speech recognition, and a companion robotic navigation unit, Koode Bot aims to alleviate heavy hospital staff workloads, streamline patient intake, and secure sensitive medical data—all without transmitting a single byte over the internet.
Main Facts
Koode Bot is an integrated, offline-first hospital kiosk and robotic escort system engineered to handle initial patient intake, intelligent department routing, and physical guidance. Built around a powerful local hardware stack, the system addresses two major pain points in modern healthcare: administrative bottlenecks and data privacy vulnerabilities.
At the core of the kiosk is a Raspberry Pi 5 paired with an 8 GB RAM configuration and a Hailo-8 AI HAT accelerator. When a patient approaches the touchscreen kiosk, they can select their preferred language—with support for Malayalam, Hindi, and English—and verbally describe their symptoms into a connected USB microphone.
[Patient Speech]
│
▼ (USB Mic)
[Raspberry Pi 5] ──► [faster-whisper (STT)] ──► [Gemma 4 E2B via Ollama (LLM)]
│ │
▼ ▼
[SQLite DB] ◄─── [JSON Clinical Report Generation] ◄────────┘
│
├─► Displays Token (e.g., K0419001) & Department on Touchscreen
└─► Publishes MQTT Message (koode/bot/navigate)
│
▼
[ESP32-S3 Robotic Escort] ──► [Ultrasonic Obstacle Avoidance / DC Motors] ──► [Guides Patient]
Rather than sending voice data to cloud-based servers, the audio is processed locally on the device using faster-whisper for speech-to-text (STT) conversion. The resulting text is then fed directly into the Gemma 4 E2B language model running locally via Ollama.
The LLM conducts a structured, multi-turn clinical interview, asking dynamic follow-up questions based on the symptoms reported. Once the interview concludes, the model generates a comprehensive JSON-formatted clinical summary containing the patient’s department assignment, urgency level, and case overview. This data is logged into a local SQLite database, and a unique token is displayed on the screen. Simultaneously, the kiosk fires an MQTT message to a companion robot powered by an ESP32-S3 microcontroller, which physically guides the patient to the correct hospital department using ultrasonic obstacle avoidance.
Chronology of Development and Workflow
The conceptualization and execution of Koode Bot represent a masterclass in localized edge computing architecture. While the exact timeline of its initial conceptualization stems from developer "lil-shan’s" open-source repository, the step-by-step operational workflow of the system unfolds in a carefully timed sequence:
- Patient Interaction & Language Selection (0–10 seconds): The patient approaches the kiosk, views the official touchscreen display, and selects their regional language preference (Malayalam, Hindi, or English).
- Symptom Collection via Voice (Variable time): The patient speaks naturally into the USB microphone. The system records the audio and passes it to the faster-whisper engine. Speech recognition typically takes between 2 to 6 seconds to accurately transcribe regional dialects and vernacular.
- Conversational Triage Interview (1 to 3 minutes): The transcribed text is routed to the Gemma 4 E2B model running locally on Ollama. The model initiates a structured clinical interview, conducting up to 12 question-and-answer exchanges to narrow down symptoms, gauge severity, and rule out immediate red flags. Individual model response latencies hover between 10 to 30 seconds.
- Report Compilation and Database Logging (60–120 seconds total): Upon interview completion, the system synthesizes the dialogue into a structured JSON report. This process—spanning loading times (~35 seconds), speech processing, and conversational turns—takes roughly 1 to 2 minutes from start to finish. The report is instantly committed to a local SQLite database.
- Token Generation and MQTT Dispatch: The kiosk displays a standardized queue token (e.g., format
K0419001, comprising an initial ‘K’, month and day, and a sequential daily counter) along with the assigned medical department. Simultaneously, a local Mosquitto broker publishes an MQTT message to the topickoode/bot/navigate. - Robotic Escorting (Real-time): The companion robot, driven by an ESP32-S3, receives the MQTT payload in approximately 1 second. Utilizing HC-SR04 ultrasonic sensors to detect obstacles and an L298 motor driver to manage dual 12V DC motors, the robot navigates hospital corridors to lead the patient directly to their designated waiting area.
Supporting Data and Technical Specifications
Deploying an LLM and multi-modal pipeline locally on edge hardware requires careful resource management. Below is a comprehensive breakdown of the hardware specifications, software stack, and operational performance metrics associated with the Koode Bot project:
Hardware Components (Kiosk & Robot)
- Kiosk Processing Unit: Raspberry Pi 5 (8 GB RAM recommended base).
- AI Acceleration: Hailo-8 AI HAT accelerator.
- Interface & Audio: Official Raspberry Pi touchscreen and a standard USB microphone.
- Robotic Unit Controller: ESP32-S3 microcontroller.
- Robot Locomotion & Sensors: HC-SR04 ultrasonic sensor (obstacle detection) and L298 motor driver powering 12 V DC motors.
Software Architecture
- AI Model: Gemma 4 E2B (quantized format:
Q4_K_M) executed via Ollama. - Speech-to-Text (STT):
faster-whisperfor lightweight, localized audio transcription. - Backend Framework: Python Flask application paired with an SQLite database.
- Messaging Protocol: MQTT via Mosquitto broker for local network communication.
- Robot Firmware: Arduino-based environment utilizing PubSubClient for MQTT handling and ArduinoJson for data parsing.
Performance Metrics & Operational Constraints
- Total Interview Duration: Handles up to 12 back-and-forth conversational exchanges per session.
- Processing Latencies:
- Model Loading Time: ~35 seconds.
- Speech Recognition (STT): 2–6 seconds.
- LLM Response Latency: 10–30 seconds per turn.
- Overall Report Generation: 60–120 seconds per patient.
- Resource Footprint: System peak RAM utilization sits at approximately 9 GB (leveraging the Raspberry Pi 5 architecture and caching).
- Network Overhead: Zero. Completely air-gapped from cloud infrastructure.
Implications for Healthcare and Community Health
The introduction of systems like Koode Bot carries profound implications for the future of localized healthcare delivery, clinical efficiency, and medical data security.
1. Radical Data Privacy and Compliance
In an era dominated by stringent data protection frameworks (such as HIPAA in the United States, GDPR in Europe, and equivalent regional data localization laws), handling patient health information (PHI) remains a high-liability enterprise. Cloud-based medical AI solutions require transmitting audio recordings and sensitive symptom descriptions across external networks to remote server farms. Koode Bot completely neutralizes this attack surface. Because every single computational process—from audio transcription to symptom analysis and database storage—happens entirely on-device, patient data never leaves the physical room. This air-gapped security model makes it an ideal deployment candidate for secure government facilities, rural clinics with strict compliance standards, and military installations.
2. Mitigating Staff Burnout and Administrative Fatigue
Hospital reception desks and triage nurses face relentless pressure, often managing long queues of patients while simultaneously performing administrative data entry, preliminary questioning, and physical guidance. By automating the initial intake interview and producing standardized clinical summaries, Koode Bot acts as an intelligent digital assistant. Doctors receive a clean, structured JSON report detailing the patient’s history, reported symptoms, and algorithmic urgency rating before the patient even enters the examination room. This drastically cuts down intake interview times for human staff, allowing physicians to focus entirely on diagnosis and treatment.
3. Epidemic Detection and Public Health Surveillance
One of the most exciting potential vectors for Koode Bot extends beyond individual clinic management. Because the system aggregates localized symptom logs within its SQLite database, networked or periodically audited arrays of these kiosks could theoretically serve as early-warning sensor arrays for regional disease outbreaks. By anonymously analyzing aggregate symptom trends (such as a sudden, localized spike in respiratory distress indicators or atypical fevers reported across multiple kiosks in a specific municipal sector), public health officials could gain real-time epidemiological insights days or weeks before traditional reporting mechanisms register a trend.
4. Overcoming Infrastructure and Connectivity Barriers
Many developing regions, remote rural outposts, and disaster-relief zones suffer from intermittent, costly, or entirely non-existent internet connectivity. Traditional smart-hospital solutions fail instantly in these environments. Koode Bot’s reliance on local edge computing proves that advanced artificial intelligence does not require a cloud umbilical cord. It democratizes access to smart medical infrastructure, bringing automated triage and robotic navigation to clinics that lack stable broadband connections.
Current Challenges and Future Outlook
Despite its impressive capabilities, the project developers note that the system is still evolving. Currently, the department misclassification rate sits at approximately 15%. While acceptable for initial pilot testing and low-acuity routing, reducing this error margin will be critical as the project matures, requiring fine-tuning of the Gemma prompt engineering and expanded clinical dataset training.
For developers, engineers, and healthcare innovators interested in exploring, testing, or deploying the system, full documentation, Raspberry Pi and ESP32 codebases, and Ollama configuration files are publicly available via the official open-source repository hosted on GitHub under shan/Koode.
As edge hardware continues to shrink in size and grow in computational capability, projects like Koode Bot signal a paradigm shift: the future of medical AI is not locked away in distant server warehouses, but operating quietly, securely, and autonomously right on the hospital floor.
