The original application could summarize a passage, generate text with GPT-2, and recognize faces by comparing facial embeddings. A chatbot called LANA added speech input and spoken responses. Its core idea was to make several model-driven tasks available through one desktop interface.
Miwl 2 retains that concept while reconsidering how the parts work together. The rebuild focuses on text, voice, and visual inputs, leaving the original stock and sales features out of scope. Local processing is the default, with an explicitly chosen cloud option for text. Reliability and control over data now guide the implementation.
The original system
The Python and Kivy source preserves the mechanisms behind the features. Summarization used a model pipeline; face recognition compared numerical representations of detected faces against stored examples. AI was already part of this software, which was built without AI coding assistance.
A surviving six-minute-forty-second recording provides a useful record of the original application in operation. It includes the terminal and desktop steps, movement between tools, processing time, successful output, and errors. The full sequence is preserved, with privacy masks over sensitive details in the version prepared for sharing.
The recording and source answer different questions. The video shows the interaction as it happened. The code exposes how the features were connected, including the facial-embedding system. Together, they provide a more useful starting point for the rebuild than a feature list alone.
Separating the interface from the work
The original source places interface callbacks, model setup, camera processing, and data access in the same application file. That makes the connections easy to trace. It also shows why clearer boundaries would help: generating text and processing camera frames involve work that should not prevent the interface from responding.
The rebuild needs to separate model and media operations from the controls that start them. A request should produce a result or an actionable error, while the interface continues to show its state. This separation also makes it possible to test model behavior independently and check the interface against controlled responses. The complete workflow needs to remain usable while models run and make recovery clear when an operation fails.
Local processing and explicit data boundaries
The current desktop application uses PySide6. Local text operations have been exercised through Ollama with Gemma 3. The voice pipeline has also been checked with Whisper transcription from an audio file and Samantha synthesis to an audio file.
Those checks establish specific parts of the local workflow. Live microphone input and audible playback still need verification. The distinction matters because reading a file successfully does not establish that an interactive conversation will handle device permissions, interruptions, and timing correctly.
Local processing is the default. An optional OpenAI provider is implemented for text requests, with confirmation required for each cloud request. Credentials are looked up only when that route is chosen, and a failed local operation does not silently fall back to the cloud. The provider’s actual connectivity remains unverified because no live cloud calls have been made.
The face gallery is stored locally and includes deletion controls. Keeping this data on the device and allowing it to be removed are concrete parts of the current implementation. Their behavior remains part of the application’s privacy testing as the vision workflow develops.
What has been established
At the 3 October 2026 development checkpoint, 136 automated tests pass, including synthetic vision coverage. Two synthetic MJPEG streams have also been exercised. That provides evidence about software behavior under controlled conditions.
Real-face recognition in the rebuilt application remains to be validated. The native vision path and recognition thresholds need testing with real inputs before accuracy claims would be meaningful. The original system’s working facial-embedding recognition is part of its history; the new implementation has to establish its own results.
Document grounding now includes FTS5 search, exact quotations and citation context. The native document import-to-search-to-citation workflow has been verified. The native document-removal check remains pending.
The next useful demonstration can follow a small set of tasks from input to result: process text, transcribe an audio file, import a document and inspect its search results and citations, and exercise the vision workflow. Each task needs evidence appropriate to the claim being made, including the failure cases someone would encounter outside a controlled test.
The original recording remains a practical reference throughout this work. It preserves both the concept and the experience of using its first implementation. The rebuild’s value will be measured in how clearly and dependably those tasks work in the next one.
Imagined by Aekam. Built and curated by Miwl.
