Collect the data that makes behavior measurable.
Mohamed A M Elansary, PhD — multi-model evaluation, uncertainty quantification, scientific data/HPC pipelines, and production agent regression evaluation.
Evaluation discipline
- Designed multi-model forecast comparisons across basins and hydroclimates.
- Validated imperfect observations before interpreting model performance.
- Reported uncertainty and regime-dependent failures instead of one score.
Production agent systems
- Builds GPT, Claude, and Gemini workflows at Vertexium.
- Maintains regression evaluation sets for retrieval, routing, and isolation behavior.
- Connects data, validation, provenance, and reporting in Python pipelines.
Proposed first contribution
Choose one user-visible agent behavior; define the slices and failure taxonomy; version an evaluation set with provenance; compare simple baselines; report uncertainty; then identify the next data worth collecting before scaling the campaign.
Honest fit boundary
Operational human-labeling management is a major minimum and a substantial stretch. I have supervised research staff, junior engineers, engineers, and contractors, but I have not hired or managed human data labelers or operated external labeling campaigns. I have not run frontier-lab post-training, RLHF, or RL/SFT recipes.
Role and location
Member of Program Staff, Data · Palo Alto, CA. I am willing to relocate with a relocation package and do not assume remote eligibility.
“expected base pay is $200,000 - $300,000 USD per year” · Official role posting