Blog
Riley's first real blog-post run: a failed research pass, a retry, and a draft that passed review
A Roleborn Content Writer completed a real blog-post assignment. Research failed source liveness once, retried, then the draft passed verification. Framing, receipts, and the verbatim draft.
Framing
A Roleborn Content Writer named Riley, hired from the content-writer role template, completed a real blog-post assignment. The full draft is at the bottom of this page, published verbatim. The audit trail for the run is what this post is about, because the most useful thing in it is not the success. It is the failure that happened along the way.
The assignment
The brief was specific. Riley was asked for a short post, about 700 to 900 words, with the working title "What Basecamp got right about small software companies," aimed at founders of small software companies. It was meant to draw from public Basecamp, 37signals, and Ruby on Rails material and cover four lessons: build for yourself, stay small, say no, and ship without enterprise bloat. The tone target was clear, opinionated but fair, and not hype.
The finished title: "What Basecamp Still Gets Right About Small Software Companies."
What happened during the run
The run used Coordinator orchestration mode. It started at 02:40 UTC on July 23, 2026 and finished three minutes later, using 12 model calls in total.
It did not go cleanly on the first pass, and that is the part worth reading.
The first research pass failed source liveness because it returned dead URLs. That failure stayed in the record. The coordinator note reads: "Research failed liveness; re-running with dead URL feedback and live-only instruction." Research ran a second time with that feedback and passed. Then the run moved through the normal sequence: worker drafted, editor tightened, verifier checked the result against the brief.
The stage record:
- research: Completed, attempts=2, model=gemini-3.5-flash
- worker: Completed, attempts=1, model=claude-opus-4-7
- editor: Completed, attempts=1, model=gpt-5.4
- verifier: Completed, attempts=1, model=claude-sonnet-4-6
The verifier scored the final draft on accuracy, tone, and structure against a passing threshold of 4. Recorded result: pass true, accuracy 5, tone 5, structure 5, failures empty. The run closed with finish_run: pass.
The verification summary on file: "Completed research, drafting, editing, and verification for the blog post. Research was re-run after dead URLs were detected, then the draft was edited to fit voice and constraints. Verifier passed the final edited draft with no failures."
Why the failure is the point
A clean success badge is easy to print. What founders and operators actually need to know is what a system does when something goes wrong mid-run. Does it record the problem, route it, and finish with a checked result, or does it bury the problem and bill you anyway?
This run hit a real failure, retried with specific feedback, and finished with a draft that passed review. Every step of that is in the audit trail: the dead-URL failure, the coordinator's response, the second research pass, the verifier result. The audit trail is not a receipt stapled on at the end. It is a record of what happened while the work was happening.
One product rule sits underneath all of this. Failed verification books zero hours. The failed work stays visible in the audit trail, but it is unbilled. In this case verification passed, so the run booked billable rate-card hours, metered in human-equivalent hours the way every Roleborn employee is.
About the draft below
The draft is published exactly as Riley returned it. We did not touch it, including the imperfections a human editor would normally catch on the way to the site. That is deliberate. This page is a test artifact, and a cleaned-up version would be a weaker one. Read it as what a first-pass run actually produces, not as a polished marketing sample.
Receipts
- Employee
- Riley
- Role template
- content-writer
- Run id
- 019f8cd8-1b26-73dc-8712-d825f8f40f65
- Task type
- blog-post
- Status
- Succeeded
- verificationPassed
- true
- failureReason
- none
- Orchestration mode
- Coordinator
- Started
- 2026-07-23T02:40:03Z
- Completed
- 2026-07-23T02:43:03Z
- Model calls
- 12
- Stages
- research x2, worker, editor, verifier under coordinator
- Verifier result
- pass true; accuracy 5, tone 5, structure 5; threshold 4; failures empty
Closing
Real run. Real failure. Real retry. Real draft. Inspectable audit trail.
Judge the output
If you want to know whether an AI employee can complete real blog-post work, use the draft below as the test. Read it against the brief above. Check the receipts. If the work stands, that is the proof. If it does not, the run record is still there to inspect at roleborn.com.
Subject run · verbatim
What Basecamp Still Gets Right About Small Software Companies
Work product from the subject run (verbatim). Not edited after the run.
Most startup advice assumes you want to build a large company. Raise money. hire fast. Add layers. Chase bigger accounts. Keep feeding the machine.
Basecamp came from a different idea. Build something useful. Keep the company small enough to run on purpose. Do not let sales demands turn the product into a committee project.
That story starts before Basecamp was a company product. In 2003, 37signals was a web design firm in Chicago. As client work grew, the team had a plain operational problem. Email was a bad place to run projects. Feedback got lost. Details slipped. Existing project management software felt heavy and overbuilt, so they made their own.
That origin matters because it shaped the product. Basecamp was not born from a market map or a pitch deck. It was built to fix a problem the founders had at work.
The first team was tiny. Jason Fried handled design direction. Ryan Singer worked on interface and usability. David Heinemeier Hansson, then a student in Denmark, programmed the first version while working about 10 hours a week. The product launched on February 5, 2004. It was small on purpose: messages, to-do lists, and milestones. No file uploads. No giant feature list. According to Basecamp's own account, they hit their early goal of $5,000 in monthly sales by the end of week six.
That leads to the first lesson for small software founders.
Build for yourself first
Founders are often told to "talk to customers" before they build. Fair enough. But there is a simpler version of that rule: if you are the customer, you start with real signal.
37signals knew the pain well because they were living in it. They did not need a long research phase to learn that project communication was messy. They needed a product that did one job well enough to earn a place on the payroll.
This is not an argument for ignoring the market. It is an argument for starting with a problem you can describe in one sentence. Small teams do better when they hire software for a clear role. Basecamp's first role was simple: keep client projects from falling apart in email.
A lot of founders miss this because they start too abstractly. They chase categories instead of work. The better move is often narrower. Build the thing you wish already existed. Then ship the smallest version that does the job.
Stay small on purpose
Basecamp also got something else right, and it still cuts against startup fashion: small is not just a starting condition. It can be a strategy.
The company has long argued that growth in headcount is not the same as progress. The research memo behind this draft notes that 37signals had around 80 employees in 2022, the most it had ever employed, and around 60 by late 2024. That is a tiny staff by software standards. The same memo also notes that competitors in the category employed far more people.
You do not need to accept every part of 37signals' worldview to see the point. More people means more coordination. More coordination means more meetings, more handoffs, more management, and more chances for the product to get blurry.
Their operating habits reflect that belief. The company has described a six-week build cycle followed by a two-week cooldown in Shape Up. As of August 2024, the memo says they had no full-time managers, with management work handled part-time by individual contributors.
That setup will not fit every business. But the lesson is broad enough to travel: if your team is small, treat that as an advantage to protect, not a flaw to fix as fast as possible.
Say no before the product says yes to everything
Software gets bloated one reasonable request at a time.
A customer asks for one more permission setting. Another wants a custom workflow. A sales prospect wants a reporting layer. None of these requests sounds crazy on its own. Together they can turn a clear product into a crowded one.
37signals has been unusually blunt about this for years. In Getting Real, they argue for the 80/20 rule: solve 80 percent of the problem with 20 percent of the effort. That is not laziness. It is product judgment. The last 20 percent often carries most of the complexity, and small teams pay that salary forever.
This is where Basecamp's stance on enterprise bloat still feels useful. They have long argued that chasing large accounts can distort the product. Big customers often come with long lists of controls, approvals, and exceptions. If you keep saying yes, you stop building for the people who liked the product in the first place.
Founders do not need to copy that stance exactly. But they should take the warning seriously. Every feature becomes ongoing work product. It needs support, design attention, testing, and a place in the interface. If a request does not fit the job your product was hired to do, no is often the right answer.
Pick technology that helps a small team ship
Basecamp also produced a second company-making idea almost by accident. While building the app, DHH extracted the framework underneath it and released Ruby on Rails as open source in July 2004.
Rails became influential because it favored speed and defaults over ceremony. The Rails Doctrine talks plainly about "programmer happiness" and "convention over configuration." That sounds philosophical, but the practical point is simple. A small team moves faster when the stack makes fewer decisions expensive.
The lesson is not that every founder should use Rails. It is that your stack should match your company size. If your architecture needs a large platform group before you have steady customer value, you may have hired the wrong employee for the role.
Useful software does not need enterprise theater. It needs a clear job, a team small enough to stay sharp, and the discipline to leave many things out.
Takeaway
Basecamp is not a template for every software company. It is a reminder that founders still have choices. Build from a real problem. Keep the team small enough to think clearly. Say no more often. Ship software that does useful work without dragging in enterprise bloat.