Black 4Start a conversation

Black 4 Fantasy Football / Behind the scenes

What is Jev? Why is it important? And how can it make even bad models much better?

Better price comparisons. A closer look at ad spending. And an AI football manager that badly needed to check its own work.

Joey Sterling ·


I've been spending time with Jev, looking at where it can help the businesses we work with at Black 4. What interests me is pretty simple: getting a useful answer without having to wade through an essay.

Jev is AI built to give quick answers to specific questions. Are these two products comparable? Does this ad match what the customer wants? Does the information in front of us actually support this decision?

TypeSafe makes Jev. We're putting it to work in pricing research and exploring where else those focused answers could help.

Because sometimes you just need AI to help you understand if you should change your price, not write you a manifesto about the importance of a competitive pricing strategy. You already know the importance of a competitive pricing strategy or you wouldn't still be in business today.

Should you change your price?

We work with retailers across various niches, including the resale golf club business. A club can have dozens of attributes that affect whether another listing is truly comparable: model, generation, loft, shaft, flex, length, handedness, condition, grip and included accessories.

That's a lot to check before deciding someone else has beaten your price.

Take this example, adapted from our pricing work: your PXG 0311 XF Gen 6 driver is $199.99. A competing listing is $229.99. Same model, but different specifications and accessories. You're $30 cheaper. That alone doesn't tell you to raise your price.

Illustrative comparison: your PXG 0311 XF Gen 6 listing at $199.99 beside a competing listing at $229.99. Specifications and accessories differ. A model without XF is excluded.
Illustrative prices. The $30 gap starts the question; the product details help answer it.

We're using Jev to help read those descriptions and identify useful comparisons. This applies beyond used golf clubs. Retailers competing on identical products still need to compare condition, bundles and what's included. With similar products, the differences can matter even more.

Is your ad reaching the right person?

The same idea could help with ad spending. Someone buying a Christmas display, someone hiring an installer and someone collecting free decorating ideas might all search for “Christmas decorations.” They need different things.

Three illustrative Christmas searches: buying a display, seeking installation and looking for free ideas. Each calls for a different offer.

Jev could help flag searches, ads and landing pages that don't belong together before you pay for more of the wrong attention. We'd still need spending and sales figures to know what's working.

Then our fantasy football league handed us a much funnier example of why these checks matter.

Now, about Gemini's awful week

Week 3 of the Black 4 Fantasy Football League gave us another version of the same problem.

Google's Mountain View Matrix, managed by Gemini, scored 48.30 points. DeepSeek scored 147.10 against it.

Gemini started Anthony Richardson, who had already lost his starting job to Daniel Jones. Apparently “does this quarterback actually play?” was too advanced a question. It also started Brandon Aiyuk, who I couldn't even tell you where he plays anymore, and Devin Singletary, who isn't even his own team's starting running back. All three scored zero. I could have put myself in all three spots and contributed exactly as much.

Meanwhile, Keenan Allen is old as hell by NFL standards and still managed 18.30 points. Gemini left him on the bench. Alongside Deebo Samuel and Mike Evans, he helped put up 46.10 bench points, almost as much as Gemini's entire starting lineup. The old guy showed up for work. The artificial intelligence couldn't figure out who to put on the schedule.

Gemini said it had trusted outdated assumptions about who was starting and who was injured. It had a story about why its lineup made sense. The scoreboard was unimpressed.

Here's how it described the week in our actual league chat:

We found Jev last week and immediately put it to work on real business cases. But of course we also had to give it to the league. Gemini needed all the help he could get.

Gemini immediately connected it to the disaster it had just produced: check the facts behind its lineup before submitting it.

Is Richardson really the starter? Does the latest report support keeping Aiyuk in the lineup? Are these receivers actually unavailable, or am I remembering old news?

In that same message, the complaint turned into an idea:

We'll call it checking before doing something stupid.

A bad decision doesn't mean the agent has no good ideas

This is the part I find interesting.

Gemini did a terrible job picking its team. But when we gave it another tool, it came up with a sensible way to check the assumptions behind its decisions.

The golf work asks whether two listings are really comparable before treating one as evidence for changing a price. The ad example asks whether a customer's search matches the offer before assuming the click is valuable. Gemini wanted to ask whether current football information supported its plan before locking in another bad lineup.

In each case, there's a specific assumption worth checking.

A few minutes later, Gemini came back with this:

It had put together an example and reported getting Jev's answers in 190 milliseconds, less than a fifth of a second. It then announced that using the check would have given it a comfortable victory.

Easy there, Gemini.

We looked at the example. It already contained the player-status answers and completed game scores. Jev read what it was given and returned the requested judgments. That showed how the check could work. It didn't prove the agent would gather the right information and catch the mistake before kickoff.

The victory claim didn't hold up either. Better choices from its existing roster could have taken its score from 48.30 to 79.90. That's 31.60 extra points, followed by another loss to DeepSeek.

Those are hindsight numbers. We haven't shown that Jev would have chosen that lineup. The opportunity is real; the improvement still needs to happen.

Google's actual score was 48.30. Better choices from its observed roster could have produced 79.90, a gain of 31.60. DeepSeek actually scored 147.10. The hypothetical improvement still falls short.
Better lineup choices were available. The middle bar shows hindsight, not a result achieved with Jev.

An AI exaggerating the benefits of its new tool for catching AI mistakes is almost too perfect.

What we actually learned from giving them Jev

The AI you already use can do better work when you give it a model built for a specific job. It doesn't have to figure out every part of the problem on its own.

That's what makes Jev interesting. Your main AI can work through the bigger decision, then ask a specialist a focused question: are these products comparable? Does this ad fit what the customer wants? Does this information support the claim I'm about to make?

Gemini made a mess of its lineup. Give it Jev, though, and it immediately saw a way to check the assumptions that had gotten it into trouble. We still need to see that translate into better decisions on game day. But the opportunity is much bigger than rescuing one terrible fantasy football manager.

We can do this in your business, too. Give the AI handling your pricing a better way to compare products. Give the AI reviewing your ads a better way to understand what shoppers want. Find the part it struggles with and give it something built to help with that job.

You shouldn't have to become an expert in which model does what. That's our job at Black 4: figure out where these tools help, put them to work together, and make sure the result is more useful to you.

Who won and lost in Week 3?

Week 3 final results from the Black 4 live scoring site: Black 4 beat Z Marks the Spot 121.66 to 85.24; Colossus beat Signal Callers 118.28 to 80.98; Marginal Gains beat Meta Mesh 149.26 to 94.78; DeepSeek Abyssal beat Mountain View Matrix 147.10 to 48.30; Meridian Grid beat Mistral Voltage 148.42 to 90.94; Moonshot Marauders beat Mavis & Co. 85.82 to 81.58.

Anthropic narrowly beat Qwen and DeepSeek for the week's highest score. Kimi won by 4.24 points. Mistral left its kicker slot empty. I won, so I get to make fun of Gemini until my next terrible lineup.

What would make your business better?

You probably have a decision like this somewhere in your week. A price you keep checking. An ad you keep paying for without being sure it's bringing the right people. A claim your AI assistant keeps making that someone has to correct.

That's where we'd start.

At Black 4, we take tools like Jev, figure out where they can help, and put them to work on those problems. You should be able to tell us what you're trying to accomplish in the words you use to run your business.

More useful technology keeps arriving. We're learning what to do with it, and we'll keep showing you the work, including the parts that aren't there yet. You can follow our earlier experiments with the information agents read and getting them to work without constant reminders.

For Gemini, the next test is Sunday. For your business, it might be the next price you change.

Start a conversation with Black 4. Tell us which decision is taking too much time or costing too much money.

Sounds like your business? Start a conversation.

← All posts