Springboards builds Flint to make LLMs less repetitive
Large language models are often presented as flexible, creative systems, but the article argues that many of them behave in surprisingly repetitive ways when asked open-ended questions. A simple example is the prompt “Give me a random number between 1 and 10”: mainstream chatbots such as ChatGPT, Claude, and Gemini frequently choose 7, then tend to follow with a small set of familiar alternatives like 3, 4, 8, or 9. The point is not that they always do this, but that their behavior is far more predictable than many users assume. For tasks where consistency matters, such as coding help or research assistance, that predictability can be useful. But for brainstorming, branding, travel ideas, or other creative work, it can create a kind of machine-generated groupthink.
The Australian startup Springboards is trying to address that limitation with a language model called Flint. According to cofounder and CEO Pip Bingemann, Flint is designed to produce a wider range of answers to open-ended prompts, even when those answers are less conventional. Bingemann frames this almost provocatively: while most model developers are trying to reduce hallucinations, Springboards says it “welcomes” them in the specific sense that it wants more unexpected, less standardized outputs. In a demo, ChatGPT and Claude both gave the expected answer of 7 to the random-number prompt, while Flint, after one ordinary attempt, produced the much less typical number 3.7916. The same pattern appeared in other prompts: when asked to name a type of car, mainstream models leaned toward Toyota or Honda, while Flint offered a Ford F-150. When asked for a New Balance running-shoe tagline, Claude and ChatGPT both produced “Run your way,” while Flint gave a different, if still imperfect, line.
The central problem is that many LLMs are not just repetitive inside a single model; they also converge with each other, giving similar answers across different companies and model families. The article connects Springboards’ pitch to a growing research concern. A paper titled “Artificial Hivemind: The Open-Ended Homogeneity of Language Models (and Beyond)” found a high degree of repetition across language models when they were asked open-ended questions. The researchers tested 25 models, including leading US systems and open-source models from China and elsewhere, and asked each of them 50 times to write a metaphor about time. Many of the 1,250 responses clustered around familiar ideas such as “time is a river” or “time is a weaver.” The paper won a best paper award at NeurIPS, signaling that this is becoming a serious topic in AI research rather than just a funny chatbot quirk.
The exact cause is still uncertain, but the article points to a plausible explanation: many modern LLMs are trained in similar ways, on similar kinds of data, and optimized for similar goals. If models are rewarded for coherent, safe, familiar, high-probability answers, they may naturally settle into the same grooves. OpenAI, responding to the issue, says that training for reliability and coherence can push models toward familiar answers, while pushing too hard for novelty may weaken reliability. It also notes that the research paper examined 2024 models that have since been updated. That response highlights the trade-off at the heart of the debate: users want models that are accurate and dependable, but they also want tools that can surprise them when the task is creative.
The article gives everyday examples of this sameness. Springboards cofounder and CTO Kieran Browne says most chat interfaces make users feel as if they are having a personal conversation, even though many people are receiving very similar outputs. For example, when asked to suggest band names, many models tend to use words such as “glass,” “neon,” “velvet,” or “static.” In the article’s own test, ChatGPT suggested names including “Glass Harbor,” “Static Empire,” “Neon Hearts,” and “Velvet Echo,” while Gemini suggested “Static Horizon.” Some of these suggestions may sound usable, but they are often generic or already taken, as shown when “Sofa Astronauts” turned out to be the name of an existing band.
Springboards’ broader product is aimed at creative professionals in advertising and marketing. Its tool lets users collect and move around text generated by several models, including ChatGPT and Claude, then combine the strongest fragments into new ideas. Flint is being positioned as another model option inside that workflow, specifically for moments when users need more variety rather than the same polished, average answer. The startup is not claiming that novelty is always better; it is betting that creative teams need controlled unpredictability as part of the process. In that context, an odd or imperfect answer can still be useful if it breaks the pattern and gives people a new direction to explore.
The broader significance is that LLM “creativity” may depend less on raw model size and more on how models are trained, tuned, and sampled. If many systems are optimized to avoid risk and choose the most statistically comfortable response, they may become less useful for tasks where originality matters. Flint’s approach suggests a different design goal: instead of treating all unusual output as a defect, it tries to make room for more diverse responses while still serving practical creative work. That may not replace mainstream general-purpose models, but it could become valuable as a complement to them, especially for users who are tired of getting the same slogans, metaphors, names, and ideas as everyone else.
Why it matters
- —It highlights a practical weakness of mainstream LLMs: they often give similar, high-probability answers when users want original ideas.
- —The issue matters for advertising, branding, naming, travel planning, and other creative tasks where sameness can reduce value.
- —Flint points to a possible new model category: AI systems tuned for variety and controlled surprise rather than only reliability.
Key facts
- Springboards, an Australian startup, has built an LLM called Flint that is designed to produce more varied responses to open-ended prompts.
- The article shows mainstream models often converge on similar answers, such as choosing 7 for random-number prompts or using familiar words in band-name suggestions.
- A NeurIPS best paper, “Artificial Hivemind,” found strong answer homogeneity across 25 language models tested on open-ended prompts.
- OpenAI says reliability and coherence training can make models converge on familiar responses, while stronger novelty can reduce reliability.
- Springboards plans to offer Flint inside a brainstorming tool used by creative professionals in advertising and marketing.
The full text is in the original source. Here we provide a brief summary and key facts.