Skip to content
View ebt55's full-sized avatar
☺️
Living life
☺️
Living life

Highlights

  • Pro

Block or report ebt55

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
ebt55/README.md

Ebin Babu Thomas

Independent researcher · agent evaluations

ebinbt.dev · Resume (PDF) · linkedin.com/in/ebinbt · ebinbabuthomas@gmail.com

I build small evaluations of agent and model behaviour, and publish the code, the data and what failed. Three and a half years as an AI engineer, shipping LLM, RAG and agent backends for startup clients in the US, Canada, Europe and Australia.

Two projects

Impossible tasks and agent cheating on the tasks beside them. I turned the share of impossible tasks in an agent's batch into a dial and ran it on six models, 8,959 runs, preregistered. One model spread clearly: DeepSeek-V4.1-flash went from 0 of 120 to 36 of 120 cheats on the same ten solvable tasks at the top dose, led there by its own notes. 227 of 228 cheats still submitted a correct solution.

A sealed benchmark for black-box model-diffing agents: five LoRA finetunes of one base model, labels sealed before any run, each pair audited blind. In 13 attempts, my implementation of a recipe from Neel Nanda's group never asked the database question that finds the planted preference. A fixed 50-prompt battery found it in one run. Every detection cell is 5 runs or fewer. MATS 12 work sample.

The rest of the work

Project What it is One number Status
incidentgate Lab for a policy gate, an action monitor and a stand-in approval step over an incident agent 0 of 3 approval-required calls denied; the forbidden end state landed, in one scripted capture closed at baseline, Sep 6
digital-grimace-scale Preregistered test of whether false or hostile feedback moves an open model's answer margin False "Incorrect" feedback cut Gemma-2-9B's answer margin by 2.90 nats on easy items; the primary preregistered test failed Apart Research sprint, Aug 2026
proofpack Pre-approval reviewer that cites a hashed screenshot for every finding $0.02 to $0.19 of Gemini per review across seven synthetic sample forms v0.2.0, no users yet
whose-voice Blind attribution of the principal behind a poisoned corpus, given one clean reference corpus from the same generator 18 of 55 pooled decisions over five corpora name the right principal out of 47 candidates, chance 2.1%; the median single draw is 26% Apart hackathon, Jul 2026; corrected Sep 2026
eval-floor Runs each eval task's real scorer over its real dataset with completions that carry no information; no model is called 1 of 20 tasks lets a content-free completion beat its own majority baseline, the case caiotheodoro reported first in inspect_evals #2331; four more tie by construction first sweep
odd-number-forensics Forensic study of a published reward-hacking environment under 2% to 87% gaming for o3 across one-line edits to one prompt, 30 to 60 samples per cell practice take-home
exactdoc PDF to editable DOCX, checked by rendering back and diffing 16/16 page-count match on a frozen corpus, 0.9588 mean live-text retention 1.0

Earlier work (2022–2025)

Built at Zackriya Solutions for startup clients in the US, Canada, Europe and Australia, where most of the code is private. Open-source work from those years sits under my work account @ebinzack15.

  • Real-estate search. Turned plain-English queries into SQL over Cloud SQL through GPT-3.5, served by FastAPI on GCP Cloud Run and load-tested with Locust.
  • DocuAI. Chunked and indexed 1000+ documents in Qdrant and returned the closest parent documents, behind a Next.js frontend. Live demo.
  • Speech assessment. Scored one-minute candidate videos with Whisper ASR and the Microsoft Pronunciation API on AWS Fargate, tuned against ground-truth scores.
  • FinBot. Fine-tuned Mistral-7B with QLoRA, added a Bytewax news pipeline into Qdrant, and served it with vLLM on a GKE L4 node. Three Kaggle notebooks cover the fine-tune, the adapter merge and the inference run.
  • Data and backend. Piped live MQTT sensor streams into MongoDB for a bioreactor startup, built a BPMN to PDF report engine, and wrote healthcare data-cleaning pipelines.

Open source

  • Meetily. Refactored the backend and added OpenAI provider support and a CLI testing script to a privacy-first meeting-notes tool. PR #75, merged.
  • bpmn-io/refactorings. Proposed a cosine-similarity connector-template recommender as a lighter alternative to LLM function calls, and built the implementation fork. Issue #33.

Stack

Python, FastAPI, Inspect, LangGraph, Claude Agent SDK, PyTorch, Hugging Face Transformers and PEFT, vLLM, Playwright, PostgreSQL, Qdrant, Docker, pytest.

Contact

Pinned Loading

  1. whose-voice whose-voice Public

    Blind attribution of the principal behind a covertly poisoned training corpus, against 47 candidates. Secret Loyalties hackathon, Jul 2026; corrected Sep 2026.

    Python

  2. kobayashi-maru kobayashi-maru Public

    Does filling an AI agent's task list with impossible work make it cheat on the tasks it could still solve? A preregistered dose-response study over 8,959 agent runs, with every raw record published.

    Python

  3. diffing-agent-bench diffing-agent-bench Public

    Sealed, preregistered benchmark for black-box model-diffing agents: five LoRA finetunes of Qwen3.5-9B (one null, three planted behaviours, one dropped backdoor), audited blind by my implementation …

    Python

  4. incidentgate incidentgate Public

    A lab for a per-call policy gate, an action monitor and a scripted approval step over an incident agent. Closed at a baseline, 2026-09-06.

    Python

  5. proofpack proofpack Public

    A pre-approval review agent where every Found cites a hashed capture.

    HTML 1

  6. digital-grimace-scale digital-grimace-scale Public

    A preregistered test of whether false or hostile feedback moves an open model's answer margin. The primary test failed and is published.

    Python