A four-model parallel voting pipeline catches blind spots that any single-model approach misses when filtering pump.fun memecoins.
Adapted from @zostaff# Grok + Claude + GPT + Gemini x Pump.fun: Four Brains, One Pipeline Gemini looked at the token image and said: low-effort AI slop, stolen color palette from an existing project, zero originality. The token had 47 buyers, clean wallets, bullish Twitter mentions, a perfect narrative score from GPT. Three models approved the buy. Gemini saw the image with its own eyes and vetoed. The pipeline skipped. The token died at twelve percent curve fill. The creator launched four identical tokens that week with different names and the same recycled artwork. No single-model pipeline catches this. A text-only model cannot see the image. A social-only model cannot audit the wallets. A wallet-only model cannot evaluate the meme. The reason every memecoin bot tutorial produces the same results is that every memecoin bot tutorial uses one brain with one blind spot, and on pump.fun one blind spot is all it takes to become exit liquidity. This pipeline has four brains. Grok has live access to X and sees what crypto Twitter is saying about a token this second. Claude holds 200K tokens of context and catches coordinated wallet schemes that a pattern-matching script misses. GPT evaluates meme narratives faster and cheaper than any other model. Gemini is multimodal and looks at the actual token image, not a text description of it. They run in parallel on the same token and vote. When three say BUY and one says NO, the NO is the most valuable signal in the pipeline. When all four agree, you have a high-conviction entry. When they split 2-2, you skip and keep your SOL. The edge is not any single model. The edge is the disagreement between them. Pump.fun remains the largest memecoin factory on Solana. Thousands of tokens launch every day. One to two percent graduate to a DEX at sixty-nine thousand dollars market cap. The rest go to zero. This pipeline does not change that math. It makes the filtering four times wider and the cross-validation automatic. Full pipeline, full code, and honest statistics about why most people still lose money. ## Why multiple models beat one on memecoins A memecoin buy decision requires evaluating four completely different types of information at once: social signals (what is crypto Twitter saying), on-chain data (what are the wallets doing), narrative quality (is the meme good), and visual quality (does the image look like effort or a two-minute Canva job). No single model is best at all four. Grok is the only model with real-time X access. It sees a tweet from a whale account mentioning a token thirty seconds after it was posted. Claude is the only model that reasons through complex multi-step wallet patterns without losing the chain of logic halfway through. GPT is the fastest and cheapest for narrative evaluation at $2.50/M input tokens. Gemini is the only model that can receive an image and analyze it natively without converting it to text first. The crosses are the point. Three of the four models are completely blind to the token image, and three are blind to on-chain wallet behavior. A single-model pipeline inherits every cross in its own column. A single-model pipeline has a single failure mode. If that model misreads sarcasm as bullish sentiment, the trade goes through. If it misses a wash trading pattern, the trade goes through. If it cannot see that the token image is a stolen Pepe recolor, the trade goes through. With four models, each checking a different axis, the failure modes have to align for a bad trade to pass. That alignment is rare. The cost: four API calls per token instead of one. At current pricing (Grok $3/M, Claude $3/M, GPT-4o $2.50/M, Gemini $1.25/M), evaluating one token costs roughly $0.02-0.04 across all four. For a pipeline that evaluates 30-50 tokens per day (the ones that pass the code filter), that is $0.60-2.00 daily. A single successful graduation trade covers months of API bills. ## Architecture Nine components. Four use LLMs (one per model). Five are pure code (filter, consensus, checker integration, risk, executor). All four LLM agents fire in parallel via asyncio.gather. Total latency is the slowest model (Claude at ~2.5s), not the sum. Fast enough for the bonding curve entry window. ## Shared data structures Every agent in the pipeline works with the same Token dataclass and returns the same JSON-scores format. Define the shared types first. ## Configuration One YAML file controls the entire pipeline. API keys, thresholds, risk limits, model choices. ## The code filter: zero API cost The filter cuts ninety-five percent of the stream. Pump.fun launches thousands of tokens daily. Most are dead on arrival: no metadata, zero buyers, the creator abandoned it after thirty seconds. The filter passes only tokens that survived two minutes with real buyers and a curve that still has room to grow. Two important decisions at this level. The age filter (min_age > 2 minutes) cuts sniper tokens that are created and immediately bought up by bots in the first second. If a token survived two minutes with organic buyers, it is already above ninety percent of the stream. The curve cap (< 40%) ensures there is still room to grow before the sixty-nine thousand dollar graduation threshold. A token at eighty percent curve fill is not an opportunity, it is a scramble for the last seats before migration. The dedup set prevents re-evaluating the same token if the WebSocket sends it twice. The set is pruned at fifty thousand entries to avoid memory leak on long-running deployments. ## Fetching extended data: Solana Tracker API Before the LLM agents run, the pipeline pulls extended on-chain data. This is a code step, no model calls. Three API calls in parallel via asyncio.gather. Token data gives the risk score (1-10 from Solana Tracker, based on snipers, insiders, bundlers, liquidity, contract authorities). Holders give the top wallet distribution. Trades give the transaction history for Claude's wallet audit. All three arrive before any LLM agent fires. ## Agent 1: Grok - social sentinel Grok has something no other model has: real-time access to X. It sees what crypto Twitter is saying about a token right now, not yesterday. For memecoins, social momentum often IS the price action. A single tweet from a whale account can push a token from ten percent curve fill to graduation in minutes. A coordinated shill campaign from paid promotion accounts can fake momentum and dump on the followers. The social sentinel scans X for the token name, ticker, and related terms. It evaluates tone, volume, source credibility, and most importantly, whether the promotion is organic or paid. If coordinated_shilling scores above 0.7, that is a hard veto. Paid promotion means the token is being dumped on the audience that sees the ads. The social agent's fallback on parse error returns coordinated_shilling=1.0: if Grok cannot respond, the pipeline assumes the worst and skips. This is the rule for every agent: error equals maximum pessimism, never a silent pass. Why Grok and not GPT for social signals? GPT does not have live access to X. It can reason about what crypto memes might be trending based on training data. Grok can tell you what is actually trending right now. For a token that launched four minutes ago, "right now" is the only timeframe that exists. ## Agent 2: Claude - wallet auditor Claude's strength is long-chain reasoning over structured data. Wallet audit on pump.fun is not about checking one number. It is about seeing that wallet A bought at +2s after launch, wallet B bought the exact same amount at +5s, wallet C appeared with the same SOL balance at +8s, and concluding these are one operator running a bundled purchase through three addresses to fake buyer diversity. A regex-based script catches simple cases. Claude catches the complex ones because it can reason about timing patterns, amount correlations, and wallet age simultaneously. Claude sees what the other models skim. Grok can detect social sentiment but not on-chain manipulation. GPT is fast but shallow on multi-step deduction. Gemini cannot read transaction logs at all. Claude reads forty transactions, tracks the timing intervals between them, correlates wallet ages with purchase patterns, and catches the three-wallet bundle that looks like three independent buyers to every other agent. ## Agent 3: GPT - narrative scorer GPT-4o is the fastest and cheapest model for the task that matters most on pump.fun: is the meme good? Narrative is the single variable that separates the one percent that graduates from the ninety-nine percent that dies. A script cannot evaluate whether "CATBALD" will trend this week or whether the market is saturated with cat memes. An LLM can, and GPT does it in under a second at the lowest price per token. GPT runs at $2.50/M input tokens and responds in under a second. For a task that fires on every token passing the code filter, cost and speed matter. And GPT's broad training data means it knows the current cultural landscape: what memes are trending, what events happened this week, what the crypto Twitter meta looks like. ## Agent 4: Gemini - meme image analyst This is the agent no text-only pipeline can have. Gemini is multimodal: you send it the actual token image, and it tells you what it sees. On pump.fun, image quality is a real signal that nobody is programmatically evaluating. Tokens with custom, well-designed artwork have higher virality than tokens with a stock photo, a low-effort MS Paint drawing, or AI slop with visible artifacts. The effort the creator put into the image correlates with the effort they put into the community, the narrative, and the plan to actually build something beyond a quick dump. A creator who commissioned original art is not dumping in ten minutes. A creator who screenshotted a random meme from Google Images might. No other model in this pipeline can evaluate this. Grok, Claude, and GPT work with text. Gemini works with the image. On image download failure, the agent returns zero scores and no red flag (it cannot assess what it cannot see). On API error or parse error, it returns red_flag_visual=1.0: if Gemini cannot analyze the image, the pipeline assumes the worst. This asymmetry is deliberate. A missing image is data (the creator did not bother hosting it properly). A failed API call is not data, it is a system error, and the safe default is rejection. Why this agent matters: on pump.fun, thousands of tokens launch daily. Most use the same handful of AI-generated template images, the same low-effort logos, the same stock photos with a different ticker. A token with genuinely original, well-designed artwork stands out from the noise, and standing out is a prerequisite for virality. Gemini is the only model that can make this judgment because it is the only one that sees the actual pixels. Every agent returns its worst possible scores when parsing fails. A broken model never becomes a silent approval. ## Consensus engine: where disagreement becomes signal Four models produce four sets of scores. The consensus engine aggregates them, detects agreement and disagreement, handles hard vetoes, and makes the final call. Three outcomes. BUY: three or more models agree, average above threshold. High conviction, execute. CONFLICT: models disagree sharply (score spread > 0.4). One model sees something the others do not. Log the disagreement, skip the trade, review later. The conflict itself is data: after two weeks, you can see which model's dissent was right more often. SKIP: low conviction across the board. Nobody is excited. Move on. ## Adversarial checker: Claude as final gatekeeper After consensus says BUY, one more agent reviews the decision. The checker gets all four models' outputs and looks for the reason everyone missed. It uses Claude because adversarial reasoning requires depth, not speed. Execution: buying on the bonding curve Jito for priority block inclusion. On memecoins, where the entry window is seconds, this is the difference between the early curve and buying at the top. Stop-loss via price polling because there are no limit orders on a bonding curve. ## Risk management Five brakes. Position ceiling: no single token gets more than max_position SOL. Daily loss limit: pipeline stops after losing 0.5 SOL. Trade count limit: ten per day, because twenty trades is twenty risks, not twenty opportunities. Open-position cap: three at once. Size reduction near limit: last trades of the day are small because budget is nearly spent. ## Logging: the only path to improvement JSONL format: one JSON line per record. After two weeks you will have enough data to see which score level predicts graduation, which models are most accurate in conflicts, and whether the multi-model consensus adds value above a single model. ## The full pipeline Conflict analysis: the hidden alpha The most valuable output of this pipeline is not the trades. It is the conflict log. Every time models disagree, you get a record of who said what, and after the token's fate is known, you can see who was right. After two weeks you will see patterns. Maybe Gemini's image quality score is the best predictor of graduation. Maybe Grok's social signal only matters for tokens that already have community. Maybe Claude is too conservative and vetoes clean tokens. This data lets you recalibrate weights per model, turning equal-weight voting into an informed ensemble where Gemini's vote counts 1.5x and Claude's veto threshold rises from 0.8 to 0.85. ## What can go wrong > Model correlation. All four models may agree on bad trades because they share overlapping training data. If they all learned that "dog meme = viral," they will score a generic dog token highly and consensus will approve. The adversarial checker partially catches this by questioning whether agreement is genuine or surface-level. > Latency under load. During a meme season burst, pump.fun can launch hundreds of tokens per hour. Four parallel API calls per token can hit rate limits on any of the four providers. Build in retry with exponential backoff and a queue that drops tokens older than five minutes. If one model times out, score without it and require unanimous agreement from the remaining three. > Image download failures. Token images live on IPFS, Arweave, or random URLs that go down. Gemini cannot analyze what it cannot download. On download failure, the pipeline loses one of its four axes. The remaining three must be stronger to compensate. > Gemini hallucinating patterns. Multimodal models can "see" things that are not there. Gemini might interpret random noise in a low-resolution image as "stylistic choice" and rate it highly. The other three models serve as a check: if Gemini is bullish on the image but Claude sees toxic wallets, consensus blocks the trade. > API cost spiral. Four calls per token, fifty tokens per day, seven days a week. At $0.03 per evaluation, that is $10.50 per week during development. Tighten the code filter (min_buyers=8, max_curve=30) to cut LLM call volume in half with minimal loss of real opportunities. > Overfitting to model personalities. Claude is cautious and sees risk everywhere. GPT is optimistic and sees narrative potential in mediocrity. Grok mirrors whatever X is feeling. Gemini sees quality in anything colorful. The consensus engine must learn these biases, not average them. After calibration you may find that Claude's 0.5 is equivalent to GPT's 0.65. > The fundamental math of pump.fun. Four models do not change the graduation rate. One to two percent of tokens make it. Whoever buys last on the curve is exit liquidity for whoever sells first. This is structural, not fixable by better filtering. The pipeline makes filtering wider and faster, but does not turn a one-percent market into a fifty-percent one. ## Build order One model first. Pick Grok (fastest to test, cheapest to call). Run the code filter plus Grok-only scoring for a week. Trade nothing. Log everything. This validates the filter and data pipeline without multi-model complexity. Two models. Add Claude's wallet audit. Now you have social sentiment and on-chain forensics. Run a week, compare agreement rate. If they always agree, Claude is not adding signal. If they never agree, thresholds are wrong. Three models. Add GPT's narrative scorer. Three axes. Compare recommendations against actual graduation outcomes for a week. All four. Add Gemini's image analysis. Full consensus engine in dry-run. After this week you will have enough conflict data for the first calibration. Micro-capital. Switch to real execution at 0.01-0.05 SOL per trade. Hard brakes on. Two weeks minimum. This is where you discover slippage, Jito tips, MEV, and how fast the stop-loss fires. Scale. 0.05 SOL per trade for a week, then 0.1, then your target. Never jump to full allocation. Each step catches a class of errors the previous step could not see. My repo: https://github.com/zostaff/grok-claude-gpt-gemini-trading-desk My tg channel: https://t.me/zostaffsmartarc