AterraAI · updated shortlist
LeanTeams|Prepared for Amanda Kam · September 2026

Updated shortlist, and the teams these five make

Five people are left. Three could lead the AI work, and Lucas could take either seat.

The two roles only make sense together, and the strongest individual is not always in the strongest pair. So this is built around choosing a pair rather than ranking five people.

1

The five

Open any card for the full assessment, videos, CV and code.

2

The teams

Seven pairs, what each covers, what it costs, what stays open.

3

Next steps

A two-week trial, and what it would settle.

The five

Rates are what you pay, all in.

Kevin Wolf

Costa Rica · 16 years · CTO and sole engineer of a live US product

The strongest submission received. He built a costing engine that knows exactly which of its own numbers to trust, and refuses to guess at the rest.

Available Now, no notice period · 20 hrs/week

10
score
$50/hr
$4,333/mo at 20 hrs/week
$2,167/mo at 10 hrs/week
Top of the listRead moreSee less
Intro video

What stood out

  • The most accurate submission received. Every figure right but one, and that one he refused to price on purpose because the rate pack says there is no matching rate for it.
  • Every number can be traced to a line in your documents. All twenty-six were checked against the original PDFs and all twenty-six matched exactly.
  • His system says when a quote is not ready to send, in plain words, with the reason. Nobody else built that.
  • It held up when the inputs were changed without warning. Moving the trip a year forward made it refuse the old prices rather than quietly reuse them.
  • He tells you where he was wrong. On video he volunteers that his own answer key was wrong and the AI corrected him.
  • Runs a live multi-tenant product single-handedly, so nothing about your architecture would be new to him.

What to weigh

  • No formal test suite. He built a strong integrity check instead, and it works, but that is one safety layer where the best submissions have two.
  • Twenty-one review items is more than a consultant will comfortably work through, even sorted by urgency.
  • Prices with no stated expiry are never checked for expiry. Narrow, but worth raising with him.
  • He reads from prepared notes on video and the delivery is imperfect.
  • Sole engineer at his current product, so he is used to working without anyone reviewing him.
10The technical challenge10/10 · the strongest submission receivedRead moreSee less
Kevin walks through his challenge

The short version: the strongest submission received. Every figure correct but one, and that one is a deliberate refusal he can defend. It survived every test thrown at it, and he tells you himself where his own answer key was wrong.

What it got right

  • Fifteen of sixteen figures correct, including both deliberate traps: the supplier email that quietly replaces two hotel rates, and the tour priced per vehicle rather than per person.
  • The sixteenth is a refusal, not a mistake. The rate pack says a Johannesburg meet and greet has no matching rate, so rather than borrow the Cape Town price he refused to quote it and raised it as blocking. That reading is arguably stricter than the answer key.
  • All seven things a human must resolve were caught, plus six more, split into seven that stop the quote being sent and fourteen to confirm.
  • Every figure carries a quote from the source document. All twenty-six were checked against the original PDFs by hand. All twenty-six matched word for word, so anyone can search the source and find the line.
  • His output says when a quote is not ready. It emits a plain not-sendable flag with the reason attached. Nobody else built anything like it.
  • He catches the chronology trap automatically. The rate pack was compiled two days after the email that overrides it, so ordinary date order points the wrong way. His system spots the contradiction and says so in the review note rather than quietly picking one.
  • He works out quantities from the price, not from the printed count. The quotation prints “2 x” conservation levy for a party of five. Rather than a rule about levies, his system notices the printed count disagrees with what the rate implies and flags it. That same logic would catch the equivalent error from a supplier nobody has seen before.

What was tested afterwards

  • Trip dates moved forward a year: it correctly refused to reuse expired prices and raised the affected lines.
  • A booked service deleted: it disappeared from the total rather than being silently billed.
  • A new service added with no matching price: flagged, unpriced, and left out of the total.
  • A room renamed so the price no longer matched: kept on the quote, unpriced, and raised for review.
  • Two figures deliberately corrupted in his output: his own integrity check caught both.

What to weigh

  • No formal test suite. He built a strong integrity checker instead, and it was verified to work, but that is one safety layer where the very best submissions have two.
  • Twenty-one review items is a lot for a consultant to work through, even sorted into blocking and confirm.
  • Prices with no stated expiry date are never checked for expiry. Narrow, and worth raising with him.
  • He spent roughly three hours against a two-hour brief, and says so.

On the walkthrough

He starts the live run at the beginning of the video and returns to it at the end, so the pipeline is genuinely running on screen rather than replayed. Six minutes against a five-minute ask, the closest of anyone. Twice he volunteers something against himself: that his own hand-written answer key was wrong and the AI corrected him, and that early on the model kept returning tidy output that was quietly missing most of the services, which he fixed by forcing it to commit to a count and refusing partial results. He reads from prepared notes and the delivery is imperfect.

9Background and fit for this role9/10 · runs a live product alone, and can start nowRead moreSee less

What he has built

He is the engineer behind a product used by bail bond agencies in Utah, working as its acting CTO. Someone calls, an AI answers, takes the details, looks up the case, scores the application and produces the paperwork. A person approves it, and only then is a payment link sent.

The part that maps onto your problem

There are around sixteen thousand of these agencies in the US and each has its own forms. His system reads a document it has never seen, works out what each field is for, and places the case information into it. That is the same shape as reading a supplier’s rate sheet you did not design, though it stops short of pulling a price out of one.

The answer that stood out

Asked where a human sits in the chain, he said the system scores and stops. It does not act, because acting on some of those grounds would be illegal. He arrived at approve-before-money-moves through a real constraint rather than a design preference, which is why I believe he would hold that line here.

What the exercise settled

Both of the open questions. He can pull structured figures out of a supplier document under time pressure, and he has a way of checking his own accuracy: he wrote an answer key by hand from your documents first, then had his program grade itself against it. When the two disagreed on one line, he investigated, found his own key was wrong, and changed it rather than the program. That is the habit you want in whoever owns the numbers.

He turned the exercise around the same night it was sent, and he is still available to start with no notice period.

Where Kevin is strong

Dashed line is what the work needs. Tick a name to compare.
Compare with

Emmanuel Abugauch

Córdoba, Argentina · 8 years · Head of AI at a payments company operating in 80 countries

He has built this system before, commercially. At Yuno, supplier terms arrive by website, PDF, email and phone call, and his team turns them into one clean catalogue that pricing runs against.

Available Part-time alongside his current role · start date to confirm

10
score
$50/hr
$4,333/mo at 20 hrs/week
$2,167/mo at 10 hrs/week
Most provenRead moreSee less
Intro video

What stood out

  • He has done this exact job before, at a company handling supplier documents from over a thousand sources.
  • The most rigorous work of anyone. His answer key provably could not have been copied from his own program's output, and he reported his own score honestly, including the parts it did badly on.
  • He found his own AI inventing a source that did not exist, and opened his submission with it rather than hiding it.
  • It behaved correctly on every test thrown at it, with no dangerous outcomes at all.
  • Runs a twelve-person AI team and still writes the hard parts himself, so he can both make the calls and build.
  • His own estimate is that he could deliver the full MVP, back end, front end and AI, in three weeks part-time.

What to weigh

  • He thinks a system like this should eventually run without a person signing off, earning trust through measurement instead. Considered, and the opposite of approve-before-sending. Worth hearing him out before he owns the architecture.
  • He works mainly in Go rather than the language your product is written in.
  • Easier to follow in writing than on video. His walkthrough shows nothing running and opens with how the system works rather than what it produced.
  • He operates at a much larger scale than yours, so there is some risk of building for a bigger problem than you have. He raises this himself.
10The technical challenge10/10 · the most rigorous submission receivedRead moreSee less
Emmanuel walks through his challenge

The short version: he found every problem hidden in the documents, his figures reconcile exactly against our answer key, and he reported a fault in his own work that nobody asked about.

What he found

Seven problems were deliberately hidden in those documents. He found all seven, including the one that matters most: a villa the client booked that has no price anywhere. His quote comes to $30,140 and he says plainly that it cannot be sent. Five services have no price at all, and one of them stops the whole thing.

The test that separated him

I changed his source data four ways without telling him: renamed a room type, deleted a service, added one that did not exist, and moved the whole trip forward a year. Each time it did the right thing. Nothing was quietly dropped, nothing was quietly charged, and when the dates moved past the end of the rate sheet it flagged every line rather than reusing last year’s prices.

The moment worth knowing about

One line in the rate sheet is damaged, so the text comes out garbled. His system worked out the right price anyway, then wrote a tidy quotation of a source that does not exist in the document. The AI got the number right and invented the evidence for it.

He found that himself, built a check that detects it, and opened his write-up with it. He is the only person in this search who has reported a fault in their own work.

One line in his video to read in context

He says “actually, my tests are wrong.” What he means is that he wrote his tests first, they exposed gaps, and he ran out of time to close all of them. The suite passes and I ran it. He is pointing at his own published accuracy report, which shows where his system scores badly. It is a disclosure rather than a defect, but it sounds worse than it is.

How he handles being unsure

Every price is sorted into one of five states: confirmed, replaced by a later email, assumed, based on an expired rate, or not found. The label is not a judgement call. It comes from six yes-or-no facts anyone can check. Every figure names the document and the row it came from, so you can verify a number without reading any code.

9Background and fit for this role9/10 · he has built this system before, commerciallyRead moreSee less

What he has already built

He leads AI at Yuno, a payments company operating across 80 countries. Payment providers each publish their terms differently. One has a decent website, one sends a PDF, one replies by email. His team built something that reads those sources, turns them into one structured catalogue, and tests each entry against the provider’s real system until it works. Adding a new provider went from about five weeks to hours.

Why that matters here

It is your problem in a different industry, and the important half is the second one: pricing never touches the original documents. It runs against the checked catalogue, offline, with no AI involved. That is the design that stops a made-up number reaching a client.

He also built the system that protects card data at Yuno, which is what allowed the company to be certified to handle payments at all. He is used to work where being wrong has consequences.

The thing he says in his video that stuck with me

“It is really hard to create a prototype and move this to production. In test you can see a PDF with a few tables in one format, and in production you get another PDF with other types of columns and other tables.” That is precisely where you are: something works, and then real supplier documents arrive.

How he comes across

His English is fluent and he is easy enough to follow, but he pauses and restarts more in conversation than in writing, and he tends to explain how something is built before saying what it produced. Both videos open with architecture rather than the result. The substance is all there, it just arrives a minute or two later than you would want.

The one thing to put to him directly

He argues a system like this should eventually run without a person signing off each time, earning that trust through measurement instead. It is a considered position with a real argument behind it, but it is the opposite of approve-before-sending, and it is worth hearing him defend it before you decide.

Where Emmanuel is strong

Dashed line is what the work needs. Tick a name to compare.
Compare with

Barbara Neves

Brazil · previously at Multimodal, laid off with her whole team

The most usable review list in the search, and the only candidate who went away and corrected herself without being asked to.

Available Immediately, no notice period · full flexibility on hours

9
score
$50/hr
$4,333/mo at 20 hrs/week
$2,167/mo at 10 hrs/week
Fastest to startRead moreSee less
Intro video

What stood out

  • Every figure correct. Her list of things needing human review is the only one a travel consultant could pick up and work through as it is, sorted by urgency with the money at stake on each line.
  • She caught something the answer key had missed. The rate pack was compiled two days after the email that overrides it, so ordinary date order points the wrong way and only the written precedence rule saves you.
  • She kept a log of the times her AI was confidently wrong, and the rule she wrote because of each one.
  • She sent an unprompted follow-up document two days after her call, asking first whether it would be overdoing it. It answers the ownership question she could not answer live.
  • No notice period at all, so she could start this week.

What to weigh

  • Her code could not be run. Everyone else's could, and running them is where most of the findings came from, so hers is the work verified least.
  • No TypeScript, which is the language your product is written in. Her follow-up document stayed deliberately neutral on this.
  • Considerably stronger in writing than in conversation. That gap is the real trade-off with her.
9The technical challenge9/10 · every figure correct, and the most usable review listRead moreSee less
Barbara walks through her challenge

The short version: she priced a real supplier rate pack correctly, and told us which of her own figures she trusted least before we asked.

What the challenge asked

Everyone received a genuine 26-line supplier rate pack, a real operational quotation and a supplier email correcting some of the prices. The task was to produce costings a system could bill against. It is a compressed version of the job: read messy documents, get the numbers right, and be clear about what is still uncertain.

How the four compared

What we testedBarbaraFabioAndreiGonzalo
Figures extracted correctlyYesMostlyYesMostly
Every number traceable to a documentYesYesYesPartly
Arithmetic kept out of the AIYesYesYesYes
Expired rate held back correctlyYesYesYesContradictory
Refused to guess a missing priceYesYesYesInconsistent
Flagged its own uncertaintyYesPartlyYesPartly
Safeguards hold when runYesYesYesSeveral fail

The moment worth knowing about

The rate pack contained two villas with nearly identical names. Her system took the price of one and applied it to the other, then presented it as a real supplier rate. On the page it looked entirely legitimate.

She caught it herself. Not by reading every line, but because she runs a review process over her own AI output that checks results against the source documents. That process flagged the line, and her final submission marks it unresolved rather than priced.

This is the single most relevant thing in her submission. The danger with automated costing is not the error you can see, it is the plausible-looking figure nobody questions. She has a habit that catches those.

The supplier email

The rate pack was compiled on the 14th. A supplier email correcting some prices was dated the 12th. She identified the conflict, recorded it as an item needing review, and used the email rate because the pack's own terms say correspondence takes precedence. Correct, and she showed her working when asked.

Where she was less strong

She built a check to catch invented prices and had not tested it against one deliberately. She has the right instinct and had not yet closed the loop on proving it works.

9Background and fit for this role9/10 · strong judgement, available immediatelyRead moreSee less

The short version: she has done a version of this job before, several times, and it went well. The two open questions are how she works alongside your product engineer and how she handles an early-stage pace.

What she has already built

For a US auto-lending credit union: a system that reads loan application documents and produces credit decisions a human can audit, connected to the client's banking system. 99% accuracy across 1,000 document packages of about 30 pages each, 47 fields per application.

For a European telecom: document classification and extraction across four languages including handwritten forms. Processing time per document dropped from six minutes to one and a half, and daily throughput tripled.

Both are your supplier rate-sheet problem in a different industry, in settings where a wrong figure gets noticed.

Why this matters for AterraAI

The risk with this kind of system is not that it breaks loudly. It is that it drifts. A supplier changes their rate sheet format, the extraction starts misreading one field, and nobody notices for two months because the numbers still look reasonable. Barbara is the candidate who talks about that scenario unprompted and has the habits to catch it.

How she works

She has led 22 customer-facing AI projects from unclear requirements to production, mentored engineers, and made product decisions on incomplete information. In her interview she described keeping written notes of her working sessions and running one AI tool's output past a different one to check it. That is a person who assumes her first answer might be wrong.

The two things to settle

Code review. Part of this role is reviewing your product engineer's work. Her knowledge of that particular language is a few years old and she would use AI tools to help. That is a reasonable approach and it is worth agreeing explicitly how the two of them will handle it, rather than assuming.

Multi-tenancy. Keeping one client's rates invisible to another is central to what you are building. Her answer on this was thinner than the rest of her interview. She understands the concept and has not built it herself.

Barbara against what this role needs

Add to chart

The dotted outline is what we think the role needs. These positions are our judgement, offered as a summary of the assessment above.

What she sent afterwards, unprompted

Two days after her call she asked whether a follow-up would be overdoing it, then sent a short document answering the things she had not answered well live: what she would own versus what an engineer would own, why a non-technical founder cannot simply prompt Claude directly, and how a rate change should propagate to a proposal a client has already seen. Nobody asked her for it. She was also candid in it that she has never designed a permissions system from scratch and would use an established pattern rather than invent one.

Where Barbara is strong

Dashed line is what the work needs. Tick a name to compare.
Compare with

Lucas Monteiro

Rio de Janeiro, Brazil · full-stack 9/10 and lead AI 8/10

The only person who can hold either seat, at the lowest rate in the group, and the only one who works natively in your stack across both roles.

Available Now · same rate for either role, part time or full time

9
score
$30/hr
$5,200/mo full time
$2,600/mo at 20 hrs/week
Best valueRead moreSee less
Intro video

What stood out

  • He scored in the strong band on both challenges. Nobody else did.
  • Every figure right on the AI challenge, and the clearest presentation of results of anyone. His output is written for someone who does not read code.
  • He built approve-before-sending into his design without being told to, including a check that catches the AI inventing figures when it explains its own work.
  • One flat rate whether it is ten hours or forty, AI or build. The role can be reshaped without renegotiating.
  • He takes correction quickly. Told his video was too long and that he had described two things inaccurately, he re-recorded the same night inside the limit and corrected both himself on camera.

What to weigh

  • The AI does not do the hardest step. Deciding which price applies to which booked item, he did by hand rather than automating. He was upfront about it in writing and it is the first thing he would finish.
  • One real fault. With the dates moved a year forward his system quietly used expired prices and called them reliable. He treated it as a feature he had not built rather than a fault.
  • He could not give an example of catching his own AI being confidently wrong. Kevin and Barbara both could.
  • Clarity when speaking is the thing to judge for yourself. Both videos are here, including the shorter second one.
8The AI challenge8/10 · every figure right, one step still done by handRead moreSee less
Lucas walks through the AI challenge

The short version: every figure right and the clearest presentation of results of anyone, but the step where the AI decides which price applies to which booked item he did by hand rather than automating. He said so in writing before anyone asked.

What it got right

  • All sixteen figures correct, including both traps.
  • All seven things needing human judgement caught, plus two more nobody was looking for: a trailer charged to a party below the minimum size, and a transfer priced off another city's rate.
  • The clearest output in the search for a non-technical reader. It leads with the subtotal it stands behind, keeps the shaky money visibly outside it, and groups what is left by what a person has to do about each one.
  • Provenance carries the replaced price too, so a line tells you the story of a rate changing rather than only the winner.
  • He built a check on the AI's own explanations. When it answers “why is this number here?” in plain English, a separate piece of code verifies every figure it cites and marks anything invented as unverified.
  • Twenty-two tests, all passing, with automated checks running on every change.

What to weigh

  • The AI does not do the hardest step. It reads the documents accurately, and a separate calculator works out the totals, but the part in between, deciding which price applies to which booked item and what to multiply it by, he typed in by hand. Several of the flags that look most impressive were written by him rather than found by the system. He listed this first among what he would finish next.
  • One real fault. With the trip dates moved forward a year, his system quietly used the old prices and marked them reliable. Asked about it, he had already checked it before the call, but treated it as a feature he had not built rather than a fault. It is the one place his system stops knowing what it does not know.
  • No examples of catching his own AI being confidently wrong. Both Kevin and Barbara had them.
  • Both the challenge and the video ran over the limits set.

On the walkthrough, and what happened after

His first video ran thirteen minutes against a five-minute ask, and in it he described two checks as automatic when they were the parts he had done by hand. His written notes had it right. Told about both, he re-recorded the same night, came in at four minutes, opened with the result rather than the architecture, and corrected both points himself on camera without being asked to. Both videos are here; the second is the shorter one.

9The full-stack challenge9/10 · every figure correct, and the only screen that surfaces a failing sectionRead moreSee less
Lucas walks through the full-stack challenge
The screen he built

The short version: every number correct, including one our own brief got wrong. His is the only screen that tells a consultant when an entire section of a trip is barely profitable.

What the challenge asked

Rebuild your pricing screen from your reference designs, using a real quotation and real supplier rates. Three situations were included that your designs do not cover, and the brief deliberately did not say what to do about them. That is where judgement shows rather than skill.

His figures against the correct answers

What we checkedCorrect answerHis screen
Trip total$24,995.46$24,995.46
Remaining budget$204.54 under$204.54 under
Gross profit$2,789.25 · 12.4%$2,789.25 · 12.4%
Accommodation section15.3% margin15.3%
Transport section1.2% margin1.2%
Activities section15.3% margin15.3%

He is the only candidate who shows the trip total to the cent rather than rounding to the nearest dollar, which let us confirm his calculations do not drift as a quote grows. One figure in our own written brief contained a typo, and he calculated it correctly rather than copying our mistake.

The trap the challenge was built around

The data contains one flight discounted below cost to hold the trip under the client's budget. The trip still looks healthy at 12.4% margin, because the loss hides inside the blend. Transport as a section is really running at 1.2%.

His screen is the only one of the three that says so. The row goes fully red, the loss is labelled on the line, a counter at the top jumps straight to it, and the real section margin is shown. His written note explains why: so the loss "can't hide behind a healthy total."

How the three compared

What we testedLucasDiegoGiancarlo
Every figure correctYesYesOne line overridden
Shows a section losing moneyYesCalculated, not shownReports it as healthy
Over budget: warns and allowsYesWarns, then blocksYes
Service below cost: flaggedYesYesPrevented instead
Service with no rate: total markedYesYesShown as complete
Edits survive a reloadWired upWired upNot attempted
Tests and documentationYesNoneNone
Matches your visual languageClearClosestSolid

Where he disagreed with your design, and was right

Your reference screen calculates gross profit in a way that counts the travel agent's commission as Aterra's profit. The written brief defines it correctly. Lucas noticed, followed the brief, built a panel that explains the difference line by line, and left a note in the code for whoever comes next:

"Pixel perfect applies to layout, typography, colour and spacing only. Never to reproducing a calculation error."

The code underneath

All three submissions were read line by line. His is the strongest. The pricing formula lives in exactly one place, so a row and a total can never disagree, which is the most common way a pricing screen ends up quoting two different numbers for the same trip. He shipped automated tests and continuous integration, the only candidate who shipped any tests at all.

See the code he wrote
9Background and fit for this role9/10 · strong work in your stack, at the lowest rateRead moreSee less

The short version: the best submission we received for this role, from the cheaper of the two serious candidates. The one thing to test yourself is how easily you follow him.

Where he works now

Senior full-stack at SmartHow, an industrial AI platform, for the past year and a half. The work is close to yours in shape: complex interfaces in the same technology your product uses, backend services and databases behind them, background processing that runs long jobs while the user watches progress, and AI features built on document processing.

Two things there matter for you. He built a cost-tracking and billing reconciliation system for AI spend that holds 0.4% variance between what was recorded and what was invoiced, which is the same discipline your costing engine needs. And he rebuilt a legacy editor into a modular architecture with 85 automated tests, delivering a full platform version in ten days. That is the instinct to improve what exists rather than rewrite it, demonstrated instead of claimed.

Before that, four years at an agency shipping more than 50 production applications, one of which lifted a client's conversion by 200%.

His introduction video is not really an introduction

The second half is an unprompted architecture proposal for AterraAI: where an AI agent should sit, what it must never be allowed to do, and the only two ways a number should be able to enter your system. His summary of it, "the agent reads, classifies and writes prose; it does not compute money and it does not approve anything," is the principle that keeps your figures trustworthy, and he arrived at it before anyone asked him to think about the problem.

He also volunteered his own gaps without being asked. He has never used two of the specific tools in your stack, though he has built the mechanisms underneath both.

What to weigh

His spoken English is fluent, strongly accented and takes real effort to follow. His written reasoning is the best of anyone we screened, so the gap between how he reads and how he sounds is unusually wide. Much of this work happens in writing, which softens it. You would still be speaking to him most weeks, and this is the one thing about him that only you can judge.

Lucas against what this role needs

Add to chart

The dotted outline is what we think the role needs. These positions are our judgement, offered as a summary of the assessment above.

Where Lucas is strong

Dashed line is what the work needs. Tick a name to compare.
Compare with

Diego Filastro

Brazil · currently at Abercrombie & Kent

The only person who built the live-updating pricing screen you designed, and the easiest of the group to talk to.

Available Reducing hours at his current role · starts with a trial

9
score
$45/hr
$7,800/mo full time
$3,900/mo at 20 hrs/week
Clearest communicatorRead moreSee less
Intro video

What stood out

  • The only candidate who built the recalculating pricing screen end to end, which was the point of that exercise.
  • He works inside a luxury travel operator right now, so he already knows how consultants actually work.
  • His pricing logic caught a subtlety most engineers miss: he weights the average markup by cost rather than taking a plain average, because a plain average would lie.
  • On the back end he described a system where some data was shared across every customer and other data locked to one, with checks so nobody could see anyone else's. He was upfront that he worked inside that system rather than designing it.
  • He says the price should always be calculated on the server and tested before the interface, because a wrong price that looks right is worse than a button that misbehaves. That is the sharpest instinct anyone showed about where the risk sits in your product.
  • He volunteered a mistake in your own mockup rather than quietly working around it.

What to weigh

  • No tests on his pricing logic. A reasonable call inside two hours, and the one thing Lucas did do in the same time.
  • Not a fit for the lead AI seat. He would be building the product, not designing the AI layer.
  • He has not designed a multi-tenant system from scratch, only worked competently inside one.
  • In a full-time role now, so this starts with a trial either way.
9The technical challenge9/10 · the same accuracy and the most finished screenRead moreSee less
Diego walks through his challenge
The screen he built

The short version: the best-looking and best-explained screen of the three, with two gaps that separate it from Lucas.

His figures against the correct answers

What we checkedCorrect answerHis screen
Trip total$24,995.46$24,995
Gross profit$2,789.25 · 12.4%$2,789 · 12.4%
Every line after repeated editsUnchangedCorrect to the cent
Per-section margins15.3% · 1.2% · 15.3%Not shown on screen

How the three compared

What we testedLucasDiegoGiancarlo
Every figure correctYesYesOne line overridden
Shows a section losing moneyYesCalculated, not shownReports it as healthy
Over budget: warns and allowsYesWarns, then blocksYes
Service below cost: flaggedYesYesPrevented instead
Service with no rate: total markedYesYesShown as complete
Edits survive a reloadWired upWired upNot attempted
Tests and documentationYesNoneNone
Matches your visual languageClearClosestSolid

The budget warning

His is the best written of the three. It names your actual clients, quantifies the overage and offers three specific fixes. It then stops the consultant from continuing. One of his three suggested fixes is to raise the agreed budget, and that is not possible anywhere in what he built.

Our view is that the software should warn loudly and let your consultant decide, because sometimes going over is deliberate. This is a design opinion rather than a mistake, and it is worth hearing his reasoning.

What he found that we did not plant

Two hotel rooms booked for the same nights, for two travellers. It was not in our answer key and nobody else noticed. He built a warning for it and used it to give the consultant a way to resolve the unpriced service.

The section margins

Worth being precise, because this is his only real gap. His pricing engine calculates the per-section figures correctly, and the code even notes that a simple average of the markups "would lie." He computes them and never displays them. It is a presentation gap rather than a misunderstanding, and a short piece of work to fix.

The code underneath

Read line by line, and clean. The pricing formula lives in one place, and nothing in the interface recalculates money independently, so a row and a total cannot disagree. The auto-save handling is the most carefully built part of his submission, covering rapid edits, failed saves, retries and network interruptions, none of which was asked for. He shipped no automated tests.

9Background and fit for this role9/10 · strongest communicator, real travel experienceRead moreSee less

The short version: a genuine alternative. If design quality and ease of communication matter more to you than $2,600 a month, he is the pick.

Where he works now

Senior front-end engineer at Bucksense since late 2023, and before that full-stack at Veckta, an early-stage startup where he worked directly with one of the founders from the beginning. Concrete results at both, and both on his CV: he reduced a page payload from around 19MB to 500KB, cutting worst-case load times from 30 seconds to 5, and built a caching layer that reduced search infrastructure costs by roughly 66%.

Why the travel experience matters

He builds the Abercrombie & Kent website, so itineraries, pricing, availability and how that information has to come together are familiar ground rather than something he would learn on your time. At Veckta he built an interactive modelling tool where changing an input immediately showed the financial impact, which is structurally the same problem as your pricing screen.

The design question

This is where he beats Lucas. His warning states are visually differentiated instead of uniform, his hierarchy is stronger, his contrast is better judged, and his screen sits closest to your existing product. He also shipped a mobile layout nobody asked for. For a founder with strong visual taste and no designer, that is not a small thing.

Nine years before software

He ran his own electrical engineering practice for nine years before switching careers. It shows in how he explains technical decisions. He is the clearest of the three and the easiest to follow.

Terms

He can start now, and asks for a two-week trial before deciding whether to leave his current role. We would suggest taking him up on it. It is a bounded commitment that answers more than an interview does.

Diego against what this role needs

Add to chart

The dotted outline is what we think the role needs. These positions are our judgement, offered as a summary of the assessment above.

Where Diego is strong

Dashed line is what the work needs. Tick a name to compare.
Compare with

The teams they make

Seven pairs, strongest first. Costs assume the lead AI seat at twenty hours a week and the product engineer full time, and drop by roughly half on the AI seat after the first two months.

Kevin + Diego

The strongest engineer in the search next to the clearest communicator. Nothing is left uncovered.

Read moreSee less
score
9
$12,100/mo

Combined coverage

Pale shapes are each person alone. The solid line is the two together.
Compare with
Lead AI seat
Kevin Wolf
20 hrs a week · Available now
Product engineer seat
Diego Filastro
Full time · Reducing hours at his current role
Cost
$12,100 a month
$2,800 a week at these hours

Strengths

  • Nothing is left uncovered. The strongest AI work in the search sits beside the strongest product surface and the clearest communicator.
  • Diego already knows the travel side, so less of your time goes on explaining the domain.
  • Kevin runs a live multi-tenant product alone, so your architecture would need no ramp.
  • Both would review each other's work from genuine strength rather than politeness.

Weaknesses, gaps and open questions

  • The most expensive option, and $2,600 a month more than Kevin and Lucas for a difference that is mostly communication and travel knowledge.
  • Neither has worked in the other's half of a codebase before, so early on you are the only person seeing the whole picture.
  • Diego is in a full-time role, so this starts with a trial regardless.
  • Uncertain: how quickly Diego can actually reduce his hours.

Kevin + Lucas

The strongest engineering pair, both in your stack, both able to start this week.

Read moreSee less
score
9
$9,500/mo

Combined coverage

Pale shapes are each person alone. The solid line is the two together.
Compare with
Lead AI seat
Kevin Wolf
20 hrs a week · Available now
Product engineer seat
Lucas Monteiro
Full time · Available now
Cost
$9,500 a month
$2,200 a week at these hours

Strengths

  • The strongest engineering pair available, and $2,600 a month cheaper than the top options.
  • Both work natively in your stack, so either can pick up the other's work.
  • Lucas can absorb AI work when Kevin's hours drop to ten a month or two in.
  • Both can start this week with no notice period on either side.

Weaknesses, gaps and open questions

  • Neither is a natural front person, so more of the translating falls to you.
  • Nobody on the team has travel experience.
  • Lucas is strong but not exceptional on product polish, and design is part of your differentiation.
  • Uncertain: how Lucas comes across to your design partners, which the trial would answer.

Emmanuel + Diego

Someone who has built this exact system commercially, with the best product surface beside him.

Read moreSee less
score
9
$12,100/mo

Combined coverage

Pale shapes are each person alone. The solid line is the two together.
Compare with
Lead AI seat
Emmanuel Abugauch
20 hrs a week · Start date to confirm
Product engineer seat
Diego Filastro
Full time · Reducing hours at his current role
Cost
$12,100 a month
$2,800 a week at these hours

Strengths

  • Someone who has built this exact architecture commercially, so the AI layer would be right the first time.
  • The most rigorous measurement discipline of anyone, which is the habit that keeps a wrong price from reaching a client.
  • Diego covers the communication and travel gaps Emmanuel has.
  • Both volunteered faults in their own work before anyone found them.

Weaknesses, gaps and open questions

  • Emmanuel's view that the system should eventually run without a person signing off would need resolving first, and it touches your founding principle.
  • Nobody here works natively in your stack. Emmanuel is Go-first and Diego is stronger on the front end.
  • Joint most expensive.
  • Uncertain: Emmanuel's start date, and how much of his week he can really give.

Lucas + Diego

Two engineers you can talk to easily, with the AI work in safe rather than exceptional hands.

Read moreSee less
score
8
$10,400/mo

Combined coverage

Pale shapes are each person alone. The solid line is the two together.
Compare with
Lead AI seat
Lucas Monteiro
20 hrs a week · Available now
Product engineer seat
Diego Filastro
Full time · Reducing hours at his current role
Cost
$10,400 a month
$2,400 a week at these hours

Strengths

  • Two engineers you can talk to easily, which matters more day to day than it looks on paper.
  • $1,700 a month cheaper than the top options, and the AI seat costs the least of any pairing.
  • Lucas covers the back end, Diego covers the product surface and the travel knowledge.
  • Both are already deep in your problem, so the ramp is shortest here.

Weaknesses, gaps and open questions

  • The AI work is in safe hands rather than exceptional ones. Lucas scored 8 where Kevin and Emmanuel scored 10.
  • Lucas leading the AI at twenty hours a week is not simultaneously your back-end engineer.
  • Neither brings the measurement discipline Kevin, Emmanuel or Barbara would.
  • Uncertain: whether Lucas closes the gap in his own work quickly once he is being paid to.

Barbara + Diego

The most usable review discipline in the search, available with no notice period at all.

Read moreSee less
score
8
$12,100/mo

Combined coverage

Pale shapes are each person alone. The solid line is the two together.
Compare with
Lead AI seat
Barbara Neves
20 hrs a week · Available immediately
Product engineer seat
Diego Filastro
Full time · Reducing hours at his current role
Cost
$12,100 a month
$2,800 a week at these hours

Strengths

  • The fastest possible start. Barbara has no notice period at all.
  • Her review list is the one your consultants would actually use, and Diego is the one who could explain it to them.
  • Both volunteered mistakes in their own work before anyone found them.
  • Diego's travel experience plus her product instinct is a strong read on what your users need.

Weaknesses, gaps and open questions

  • Neither works in your stack. Barbara has no TypeScript and Diego is stronger on the front end than the back.
  • Barbara reviewing Diego's work weekly would be a conceptual review rather than a line-by-line one.
  • Joint most expensive, and it buys the least coverage of your actual codebase.
  • Uncertain: her code could not be run, so hers is the work verified least.

Barbara + Lucas

Every technical area covered twice over, at the lowest price.

Read moreSee less
score
8
$9,500/mo

Combined coverage

Pale shapes are each person alone. The solid line is the two together.
Compare with
Lead AI seat
Barbara Neves
20 hrs a week · Available immediately
Product engineer seat
Lucas Monteiro
Full time · Available now
Cost
$9,500 a month
$2,200 a week at these hours

Strengths

  • Every technical area covered twice over, at the lowest price.
  • Barbara's review discipline plus Lucas's fluency in your stack is close to what the lead AI role actually calls for.
  • Both available immediately with no notice period.
  • Lucas covers the TypeScript gap Barbara has, precisely.

Weaknesses, gaps and open questions

  • The least comfortable founder-facing surface of any pairing.
  • No product design strength and no travel experience.
  • Barbara's work is the one verified least, because her code could not be run.
  • Uncertain: how two people who are both stronger in writing than in conversation would run your weekly rhythm.

Emmanuel + Lucas

The highest technical ceiling on paper, and the hardest team to get a plain answer from.

Read moreSee less
score
8
$9,500/mo

Combined coverage

Pale shapes are each person alone. The solid line is the two together.
Compare with
Lead AI seat
Emmanuel Abugauch
20 hrs a week · Start date to confirm
Product engineer seat
Lucas Monteiro
Full time · Available now
Cost
$9,500 a month
$2,200 a week at these hours

Strengths

  • The highest technical ceiling on paper, at the lowest price.
  • Lucas covers exactly the stack gap Emmanuel has, and can carry AI work when Emmanuel's hours drop.
  • Emmanuel's measurement discipline would raise the bar for the whole team.
  • Both are comfortable owning decisions rather than waiting to be told.

Weaknesses, gaps and open questions

  • The hardest team to get a plain answer from. Neither is a natural technical voice for you.
  • Emmanuel's position on running without human sign-off is unresolved, and this is the pairing where it matters most.
  • No travel experience and no product design strength.
  • Uncertain: Emmanuel's availability, which is the least settled thing in this document.

All seven, side by side

The three strongest are shown to start. Tick any of the others to add them, or untick to clear.

Teams on one chart

Outline only. Dashed line is what the work needs.
Showing
TeamScoreMonthlyThe main reason to pick it
Kevin + Diego9/10$12,100/moNothing is left uncovered. The strongest AI work in the search sits beside the strongest product surface and the clearest communicator.
Kevin + Lucas9/10$9,500/moThe strongest engineering pair available, and $2,600 a month cheaper than the top options.
Emmanuel + Diego9/10$12,100/moSomeone who has built this exact architecture commercially, so the AI layer would be right the first time.
Lucas + Diego8/10$10,400/moTwo engineers you can talk to easily, which matters more day to day than it looks on paper.
Barbara + Diego8/10$12,100/moThe fastest possible start. Barbara has no notice period at all.
Barbara + Lucas8/10$9,500/moEvery technical area covered twice over, at the lowest price.
Emmanuel + Lucas8/10$9,500/moThe highest technical ceiling on paper, at the lowest price.

Suggested next steps

Trial three people for two weeks, then decide with your lead

The two weeks do two jobs at once. Your lead splits the work so the most gets built, and forms a view on both engineers by reviewing what each of them produces. You come out with a fortnight of real progress and a far better basis for the decision than another round of interviews would give you.

  1. Pick the lead. Kevin, Emmanuel or Barbara. From day one they set the architecture, break the work into pieces that can run in parallel, and decide who takes what. That is also how they end up with an informed view of both engineers: they are reviewing the output rather than watching from the side.
  2. Two weeks, three people. Diego can work full time alongside his current role, or part time if you would rather judge both engineers on equal hours. Lucas is open to either. The lead stays part time until the architecture is settled and the foundations of the AI layer are built.
     Both engineers part timeDiego full time
    Diego$600/wk$1,800/wk
    Lucas$600/wk$600/wk
    Lead, 20 hrs$1,000/wk$1,000/wk
    Two weeks$4,400$6,800
  3. Give them work that is already on your roadmap. It should be worth keeping whatever you decide, and demanding enough that how each of them handles it tells you something.
  4. Decide with your lead. They will have reviewed both engineers' work and can tell you what they saw. The call is yours.