Can linear-attention RNNs recognise formal regular languages?
Attention-free recurrent models like RWKV promise linear-time inference. The open question is capacity: does the compressed recurrent state actually retain what a Transformer's full attention window retains? I built a controlled setting where the answer is checkable rather than vibes-based.
Three things worth reading in full
Where the work happened
AI Research Intern · NVIDIA
Non-Transformer sequence architectures for formal-language expressivity. Custom CUDA C++ WKV operator; synthetic L1-L4 benchmarks with 1-3 edit-distance hard negatives.
AI Engineer Intern · Meeting.ai
ASR concurrency stress-testing, speaker-verification distillation, and an active-learning correction loop for clustering ambiguity.
AI & Backend Engineer · Intergalactic Science Kingdom
Marie Chan: Go + Django backends, OpenRouter multi-model LLM routing, BDD/TDD around agentic workflows.
IT Division Coordinator · BEM Fakultas Ilmu Komputer UI
Directed development workflows and reviewed architecture across university-wide IT projects.
Engineer first, researcher by habit
I like problems where the research question and the deployment constraint are the same problem: can this architecture represent the thing at all, and can it run fast enough to matter. That is why the NVIDIA work involved writing a CUDA kernel, and why the Meeting.ai work involved a latency harness rather than a notebook.
Outside that: five semesters as a teaching assistant at UI, a certified peer counselor at Curhat Sama Panda, and enough flight-sim hours to be opinionated about approach plates.
Grouped by reach-for frequency
daily
shipped to production
familiar
Hiring in AI research, ML engineering or data science?
I reply within a day. Happy to walk through the NVIDIA benchmark design or the distillation setup on a call.

