# Renan Butkeraites — complete site corpus Generated 2026-08-08 from https://butkeraites.com. Every page of the site, in markdown, in one file. --- # Renan Butkeraites **Engineering leader · Optimization specialist** · Kennesaw, GA, US · Remote, global Engineering leader with a PhD in Operations Research and 7+ years building Python backends, AI/LLM products, and decision-support systems — from mathematical model to production. - Email: rbritobut@gmail.com - GitHub: https://github.com/butkeraites - LinkedIn: https://www.linkedin.com/in/rbritobut - Google Scholar: https://scholar.google.com/citations?user=Zh48TgYAAAAJ - ORCID: 0000-0002-5830-0466 ## Selected work ### Vehicle routing engine for Brazil's largest sugar producer *Nitryx → Progress Rail · A Caterpillar Company* Designed and led a Python optimization engine for fleet routing across 40k+ locations — started at Nitryx, carried through Progress Rail's acquisition, and taken from concept to production-ready in 16 months with a team of 6. **Outcome:** −20% fleet size · $60K/month cloud savings via in-house Distance Matrix APIs ### Client-facing exports platform for transit operators on 3 continents *Optibus · Public Transit SaaS* Led a 5-developer team owning report generation and integrations for clients across North America, Latin America, and Western Europe — including real-time transportation map feeds. **Outcome:** −50% new bugs · 2× team development velocity · key contract secured ### SIROM — a robust optimization method under uncertainty *PhD Research · UNIFESP + Polytechnique Montréal* Created a sampling-based multi-objective iterative method for optimization under uncertainty, plus "Hardness", a new robustness measure. Co-authored with Michel Gendreau. **Outcome:** Outperformed or matched literature methods in 92.5% of Bandwidth Packing cases ### Behavioral models that moved acquisition and engagement *Banco Safra · Banking* Built a Markov-chain model of app user behavior to retarget communication, and executive-level analyses that pivoted credit policy and marketing. Led fraud-pattern analysis on PIX instant payments. **Outcome:** +15% app activation · +45% high-quality credit card leads ### From 5-day to 5-minute model deployments *Porto Seguro · Insurance* Started the company's movement toward API-based models and AWS cloud computing, and shipped a freight-size estimation heuristic used in logistics planning. **Outcome:** Deployment time: 5 days → 5 minutes · 20% logistics optimization ### Forecast UTI — ICU bed forecasting during the pandemic *Public Health · COVID-19* Co-built an application forecasting intensive-care bed demand for Brazilian health services during COVID-19, published in a national epidemiology journal. Also open-sourced a classroom-occupancy optimizer for safe school distancing. **Outcome:** Published in Epidemiologia e Serviços de Saúde (2020) ## Experience ### Information Technology Manager · SG Global Group *2026 — present* Leading IT operations and technology strategy for SG LAW LLP (Kennesaw, GA): LLM-powered legal solutions for research, contract review, and document analysis; governance of confidential legal datasets (anonymization, redaction, privacy compliance); cybersecurity across AI systems; firm-wide AI adoption and training. ### Director of Technology · Fresh Codes *2025 — present* End-to-end backend engineering for AI-driven product development: distributed cloud architectures (AWS, Azure), high-throughput data pipelines, LLM-powered ML models. Mentoring a fully remote team of 4+ engineers. ### Technical Lead · Serendipe Institute of Science & Technology *2025 — 2026* Drove the institute's technical strategy and R&D execution; project prospecting, technical hiring, and mentoring of engineers and researchers. ### Founder & CEO · BSI — AI & Optimization Consulting (closed) *2024 — 2026* Founded and led a consulting firm in Campinas, Brazil, crafting bespoke AI strategies and optimization solutions for clients; operations closed in January 2026. ### Integration Consultant, Backend Python · hotglue *2024 — 2025* Designed and implemented custom data-integration solutions on hotglue's embedded ETL platform. ### Backend Python Team Lead · Optibus *2022 — 2024* Led the customer-oriented R&D team for specialized exports; pair-programming culture, knowledge-sharing initiatives ("Optifridays"), technical interviewing during rapid company growth. ### Researcher, Optimization · Nitryx Consulting → Progress Rail (Caterpillar) *2021 — 2022* Started the vehicle-routing optimization engine at Nitryx and carried it into Progress Rail after the acquisition — leading a team of 6 from conception to production-ready in 16 months, with Distance Matrix APIs and six real-time integrations. ### Data Science Specialist · Banco Safra *2020 — 2021* Mixed Media Models, Markov-chain behavior models, PIX fraud-pattern detection, executive-level analytics. ### Senior Data Scientist · Porto Seguro *2019 — 2020* API-based model serving on AWS; optimization heuristics for logistics; ML with XGBoost, scikit-learn, Gurobi/CPLEX. ### Mathematical Analyst · UniSoma *2018 — 2019* Column-generation scheduling with CPLEX for railroad workforce planning; real-time Big Data (Kafka, PySpark) NLP fraud detection. ### Research Intern · CIRRELT — Polytechnique Montréal *2018* Optimization under uncertainty, supervised by Prof. Michel Gendreau. ### PhD Candidate, Operations Research · UNIFESP / ITA (CNPq fellow) *2016 — 2021* Thesis: "Optimization under uncertainty: a new computational method, a new robustness measure, and applications." ## Publications - Renan Brito Cano Butkeraites, Luiz Leduino de Salles Neto, Michel Gendreau, "A sampling-based multi-objective iterative robust optimization method for the Bandwidth Packing Problem", Expert Systems with Applications, vol. 203, 117337, 2022. doi:10.1016/j.eswa.2022.117337 - Luiz Leduino de Salles Neto, Renan Brito Cano Butkeraites, Martins, Chaves, Horacio Hideki Yanasse, "Forecast UTI: an application for forecasting intensive care unit beds during the COVID-19 pandemic", Epidemiologia e Serviços de Saúde, 2020. - Renan Brito Cano Butkeraites, José Luiz Chela, Luiz Leduino de Salles Neto, "Efficient frontier of credit risk using Monte Carlo simulation", International Journal of Business Intelligence and Systems Engineering, 2019. - Renan Brito Cano Butkeraites, "Optimization under uncertainty: a new computational method, a new robustness measure, and applications", UNIFESP, 2021. (Advisors: L. L. Salles Neto & W. A. Lodwick) ## Open work - **SIROM**: A sampling-based method that hands you a Pareto frontier of robust solutions and lets you choose the trade-off after seeing the options, instead of committing to an uncertainty budget before you know what it costs. — https://github.com/butkeraites/sirom - **Hardness**: A continuous, normalized robustness measure for optimization under interval uncertainty — one number in [0,1] that ranks feasible solutions by how well they resist the uncertainty around them, with an exact closed form and a Monte-Carlo estimator that agree. — https://github.com/butkeraites/hardness - **The N=32 Wall**: A verified database of Costas arrays for orders 2–100, algebraic generators built over finite fields, and CP/SAT/LP experiments against the smallest orders where nobody has ever found one — or proved none exists. — https://github.com/butkeraites/costas-array ## About I started as a **mechatronics technician** maintaining flight simulators, and ended up with a **PhD in Operations Research** — because I kept asking how decisions could be made better under uncertainty. Along the way I did research with Michel Gendreau in Montréal, published the SIROM method, and created "Hardness", a new way to measure how robust a solution really is. Since then I've applied that lens in insurance, banking, rail logistics, public transit, and data platforms: always the same job, **turning a messy real-world decision into a model, and the model into production software**. I've led teams through that whole arc — hiring, mentoring, pair programming, and shipping. Today I manage technology for **SG Global Group** in Kennesaw, Georgia — leading legal AI and LLM initiatives for SG LAW LLP — while directing technology at **Fresh Codes**, an AI-driven product studio, with remote teams and global clients. Off the clock: mathematical puzzles, good food, and music — sometimes all three while putting my daughter to sleep. --- # Projects Work by Renan Butkeraites. Projects marked *demo* have an interactive version that runs in your browser. ## Interactive demos - [SIROM](https://butkeraites.com/projects/sirom/) — A sampling-based method that hands you a Pareto frontier of robust solutions and lets you choose the trade-off after seeing the options, instead of committing to an uncertainty budget before you know what it costs. - [Hardness](https://butkeraites.com/projects/hardness/) — A continuous, normalized robustness measure for optimization under interval uncertainty — one number in [0,1] that ranks feasible solutions by how well they resist the uncertainty around them, with an exact closed form and a Monte-Carlo estimator that agree. ## Case studies - [The N=32 Wall](https://butkeraites.com/projects/costas-arrays/) — A verified database of Costas arrays for orders 2–100, algebraic generators built over finite fields, and CP/SAT/LP experiments against the smallest orders where nobody has ever found one — or proved none exists. --- # SIROM > Robust optimization without having to pick an uncertainty budget first. A sampling-based method that hands you a Pareto frontier of robust solutions and lets you choose the trade-off after seeing the options, instead of committing to an uncertainty budget before you know what it costs. - **Source:** https://github.com/butkeraites/sirom - **Licence:** MIT - **Techniques:** Robust optimization, Monte Carlo sampling, Linear programming, Clustering - **Stack:** Python, OR-Tools, scikit-learn, NumPy - **Interactive demo:** runs entirely in the browser, no server ## Headline results | Measure | Value | | --- | --- | | Bandwidth Packing cases matched or beaten | 92.5% | | Pipeline runtime, after optimization | 30.034s → 1.027s (29.2×) | | Frontier envelope reproduced across constraint structures | 18 / 18 | ## The problem with asking first Plenty of optimization problems have coefficients nobody knows exactly. Demand, cost, capacity — they live in a *range* rather than at a point. Classical robust optimization handles this by making you specify, up front, how much uncertainty you want protection against. That is the uncertainty budget, and it is a strange thing to ask of a decision-maker: choose your insurance premium before anyone has told you what the policy covers. The honest answer is usually "it depends what it costs" — and you cannot know what it costs until you have solved the problem. ## What SIROM does instead SIROM takes a linear program whose coefficients are intervals and inverts the question. It **samples** the uncertain region, **solves** each realization, **clusters** the resulting solutions by how they behave, and returns a small **Pareto frontier** of candidates — each labelled with its objective value and the probability it stays feasible under uncertainty. You pick the trade-off afterwards, looking at real numbers. The method was published in *Expert Systems with Applications* in 2022, co-authored with **Michel Gendreau**, and evaluated on the Bandwidth Packing Problem, where it matched or outperformed the methods in the literature in **92.5%** of cases. ## What the frontier actually shows you Each point is a candidate plan. Moving right buys you a higher probability of staying feasible; moving up costs you objective value. The interesting part is that the frontier is rarely smooth — it has cliffs, and the cliffs are where the real decisions live. Push the target return high enough and the frontier does not just shift, it *collapses*: the set of achievable robust solutions shrinks, sampled futures start coming back infeasible, and the method tells you plainly that what you asked for is not available at any robustness worth having. ## Engineering notes The interesting work was not the maths. Getting the pipeline from **30.034 seconds to 1.027 seconds** — a 29.2× speed-up — required per-phase attribution rather than guessing, and the profile was counter-intuitive: the clustering step cost more than every linear program in the sweep combined. That result generalises. When a pipeline mixes a "real" solver with ordinary data manipulation, intuition consistently blames the solver. ## A note on fidelity A browser port of this method cannot reproduce the Python output bit-for-bit — k-means initialisation is randomised, and no two implementations will cluster identically. The defensible claim, which the repository's own benchmarks support across 18 of 18 constraint structures, is that the frontier's **envelope** is invariant: the range of objective values and feasibility probabilities matches, even when the specific set of points does not. That distinction is worth stating out loud rather than hiding. A demo that claims more than it can prove is worse than no demo. --- Renan Butkeraites · https://butkeraites.com/projects/sirom/ --- # Hardness > How much punishment can this decision take before it breaks? A continuous, normalized robustness measure for optimization under interval uncertainty — one number in [0,1] that ranks feasible solutions by how well they resist the uncertainty around them, with an exact closed form and a Monte-Carlo estimator that agree. - **Source:** https://github.com/butkeraites/hardness - **Licence:** MIT - **Techniques:** Robust optimization, Interval uncertainty, Monte Carlo, Formal verification - **Stack:** Python, Agda, FastAPI, NumPy - **Interactive demo:** runs entirely in the browser, no server ## Headline results | Measure | Value | | --- | --- | | Measure range | η ∈ [0, 1], continuous | | Closed-form evaluation | 1.6 µs | | 20,000-scenario simulation | 9 ms | | Core theory | machine-checked in Agda | ## The question nobody asks precisely "Is this solution robust?" is usually answered with a yes or a no, which is a strange way to talk about a continuous property. Two plans can both be feasible under nominal data and both survive the worst case, and still be nothing alike in how they behave in between. Hardness gives that middle ground a number. For an optimization problem whose coefficients live in intervals rather than at points, it aggregates per-constraint worst-case violation into a single scalar **η ∈ [0, 1]**, so feasible solutions can be *ranked* rather than merely accepted. ## Two engines that have to agree The measure is computed two ways, and the fact that they agree is the point. **Closed form** evaluates a generalized Irwin–Hall CDF exactly. It is the right tool for small problems and it is fast — microseconds per evaluation. **Monte Carlo** estimates the same quantity from a violations matrix, with bootstrap confidence intervals. It scales where the closed form does not. When an exact method and a sampling method built from different mathematics land on the same curve, you have evidence the definition is sound rather than an artifact of one implementation. ## Where it gets interesting Take a fixed budget — say a renewable generation portfolio with a hard $50M cap — and sweep one decision variable. Cost is constant by construction, so it cannot rank anything. Three different robustness measures then proceed to *disagree* about which plan to buy. That disagreement is the whole argument for caring about the definition. One measure slopes monotonically the wrong way. Another is pinned flat at zero across half the range, unable to express an opinion at all. η picks a plan whose realized shortfall, under twenty thousand simulated futures, is a small fraction of what the naive choice delivers — at identical spend. ## The part I am most attached to The core theory is **formalized in Agda** and machine-checked. Not because a referee asked, but because a robustness measure that is itself unproven is a peculiar thing to publish. That formalization is the audit trail: it says the properties claimed for η are not claimed on the strength of my careful reading of my own proof. ## Composition Hardness speaks the same problem schema as [SIROM](/projects/sirom/) — `lb_A, ub_A, lb_b, ub_b, integer_variables` — deliberately. SIROM *produces* robust solutions; this *scores* them. Two peer-reviewed methods that compose without an adapter. --- Renan Butkeraites · https://butkeraites.com/projects/hardness/ --- # The N=32 Wall > Two numbers nobody can answer, and a search that shows you why. A verified database of Costas arrays for orders 2–100, algebraic generators built over finite fields, and CP/SAT/LP experiments against the smallest orders where nobody has ever found one — or proved none exists. - **Source:** https://github.com/butkeraites/costas-array - **Licence:** MIT - **Techniques:** Constraint programming, SAT solving, Exhaustive search, Finite fields - **Stack:** Python, C++, OR-Tools, kissat ## Headline results | Measure | Value | | --- | --- | | Arrays verified on every CI run | 9,217 | | Orders with no known array | 18, smallest is 32 | | Exhaustive search, order 15 | 2.2 s · 743k nodes | | Exhaustive search, order 17 | does not finish | ## A permutation with an unreasonable property A Costas array is a permutation of `1..N` where **every displacement vector between two dots occurs exactly once**. Draw the dots on a grid, connect any two, and no other pair anywhere in the array has that same offset. The property is easy to state and brutal to satisfy. It is also useful: the autocorrelation of such an arrangement is nearly ideal, which is why they show up in radar and sonar waveform design. ## The wall Costas arrays are known for most small orders. Two are missing. **Nobody has ever found a Costas array of order 32 or 33, and nobody has proved one does not exist.** Sixteen more orders up to 100 are in the same state. The algebraic constructions — Welch, Golomb, Lempel, all of them built from primitive elements of finite fields — produce arrays only for orders tied to primes and prime powers. Every other order has to be searched. And search runs into a wall you can feel: | Order | Exhaustive search | |---|---| | 13 | 0.03 s | | 14 | 0.4 s | | 15 | 2.2 s, 743k nodes | | 16 | ~19 s | | 17 and up | does not finish | Two orders past a two-second problem, the same program will outlive you. ## What is in the repository **A verified database.** 9,217 arrays across 82 orders. Every one of them is re-checked against the definition on every CI run — the difference property is re-derived from scratch, not trusted. The eighteen orders with no known array are checked too, against the published table, so the repository cannot quietly disagree with the literature. **Generators, not just data.** `costas_generate.py` builds arrays from the published mathematics with no external input: exhaustive backtracking, the Welch construction, and Golomb/Lempel over GF(p^k) — including its own finite-field arithmetic, because the prime-power orders need a real field, not modular arithmetic. It reproduces the database exactly for orders 2 through 7. Above that it reproduces a subset, and the repository says so precisely rather than implying otherwise. **Search experiments.** Propagator ablations, symmetry breaking over the dihedral group, LP relaxations sharded by column window, mined clique cuts and forbidden patterns, and a SAT encoding whose size at order 32 is itself the story: **1,016,863 variables and 3,048,590 clauses**. ## What the measurements taught me The honest results are the interesting ones, and they were not what I expected. Turning the matching propagator *off* cut node counts by only about 16% at order 14 — while making order 13 roughly **four times faster in wall-clock**. The dyadic filters changed node counts by **exactly zero** on satisfiable instances. The propagators are correct and carefully built. Measured against the thing that actually matters, most of them do not pay for themselves here. That is worth more to me than another green benchmark, because it is the kind of result you only get if you instrument honestly and then publish what came back. --- Renan Butkeraites · https://butkeraites.com/projects/costas-arrays/ ---