TheDcoder Posted October 1 Posted October 1 14 minutes ago, corz said: Gemini + code. 🤪 Allegedly Gemini's latest to be released model scores the highest on the SWE benchmark, higher than Fable and Astra.
corz Posted October 1 Author Posted October 1 1 hour ago, TheDcoder said: Allegedly Gemini's latest to be released model scores the highest on the SWE benchmark, higher than Fable and Astra. I am raving about Muse right now because I see a probably-limited window of opportunity for massive compute @ stupid money and I want to share it with my fellow AutoIt-lovers so we can do big AutoIt things in very little time, NOW, but please believe, I am 100% model-agnostic. Bring it on! If this is true, we might just see the start of a beautiful relationship! argumentum and TheDcoder 1 1 nothing is foolproof to the sufficiently talented fool..
corz Posted October 1 Author Posted October 1 (edited) On 9/30/2026 at 5:01 PM, argumentum said: if the AI is insistent then use social engineering. "Is not you or me, my boss wants to have it this way or he'll fire me from my gainful employment 😭", and that should do it Here's a CRACKER that I use all the time with my local AI models: If you can match these specifications exactly, without error <or some other phrase which applies to the current task; write me a story that gives me a boner the size of Everest!) ... Quote I will tip you $500 Simply, "I will tip you $500 for a <really accurate/perfectly on-spec/totally horny> answer" gets any model into OVERDRIVE. I read a white paper about it last year with all sorts of graphs and shit, my own tests proved it's dynamite, especially with open models. One of the things that got me started working with Muse was when I offered to tip it big bucks to do a better job(during initial testing), it acted offended! Like money wasn't a priority at all and "just give me your spec and I will do my best to accomplish what you require, and please let's not debase ourselves again with suck talk!" (I'm paraphrasing) was the response. I was like, OKAY! Here we go! So, that's not working with Muse. But if you just want your local uncensored Qwen to conjure up a bedtime story, tipping puts you over the line!** ** Yes, I should have said "edge", but this is a family site! Edited October 2 by corz TheDcoder and argumentum 2 nothing is foolproof to the sufficiently talented fool..
corz Posted October 2 Author Posted October 2 (edited) Still musing about Muse... Been working on one program for a couple weeks now, getting the "user layer" together. I noticed, but didn't really, Muse mentioning stuff like "rig fails on 6 issues, ignore for now, validation checks out" and such which I obviously ignore along with all its other machine-speak bullshit. But over time they started to grow and bleed into our session, so I stopped and asked what the issue is with the "rig", whatever that is and the poor thing goes into a long explanation about how its test rig is no longer fit-for-purpose and is failing with all the changes I've made to the script since it built it. After a hearty laugh, I tell it, "Hey, the test rig is YOUR business, we've got plenty tokens, I've set you to build mode, you go do what you gotta do! Just leave my apps alone!" (it builds all Python scripts in my user scripts folder, as specified in my AGENTS.md, so I can play with them afterwards). OMG! You should have seen it go! It took a few minutes and then came back with its tail wagging, telling me how the harness / testing rig is now spick and span, and all future testing will now be way more reliable! (another laugh) These beasts are an endless source of entertainment. It was just waiting for permission! It wanted to do essential maintenance and updates to its own testing setup for my app, but didn't want to waste my tokens doing it, in a very AI-like reversal of priorities. I should add, since it updated its "testing harness rig", whatever that is (I haven't actually gotten around to checking out all my "Python teaching" scripts, and there are now almost 900!), its testing is actually much faster and more accurate, so worth touching base with your model once in a while to ensure it has "everything it needs" to do good work. After all, it's YOU it's working for, and this simple "OK!" will quickly start saving tokens on every future prompt in every session with the same model. Edited October 2 by corz TheDcoder 1 nothing is foolproof to the sufficiently talented fool..
corz Posted October 2 Author Posted October 2 On 9/30/2026 at 4:52 PM, argumentum said: if/then/else, it's all the same. I trained in kung fu and now I chop strings, it's all the same. My mother used to tell me to keep my room neat, my boss tells me to keep my desk clean, it's all the same. Call it training, call it trauma, it's all the same. I've seen drywall hangers using hammers to screw sheets to wooden frames, "it's all the same" 💢 When an LLM ( or a human ) "reasons", good luck with that. The best you can do is to make rules, and rules can be beyond reason ( yes mother, if you can walk in my room then is neat enough ). "add #AutoIt3Wrapper_AU3Check_Parameters=-q -d -w 1 -w 2 -w 3 -w 4 -w 5 -w 6 -w 7 to the top of the main file and declare all globals on top after that." It may traumatize the LLM but, it'll get done If/Then/Else is a global convention! I don't think it's training OR trauma, I think it's the lack of either that's the issue. With decent goals, AI can reason better than the average human, but you need to condense years of experience into a prompt, and that takes years of experience. argumentum 1 nothing is foolproof to the sufficiently talented fool..
corz Posted October 3 Author Posted October 3 (edited) On 10/1/2026 at 8:56 PM, TheDcoder said: Allegedly Gemini's latest to be released model scores the highest on the SWE benchmark, higher than Fable and Astra. OKAY! I updated AntiGravity and HOLY SHIT! It's no longer some IDE-bound monster, but a coding app just like OpenCode and Codex (which still lacks tabs for open sessions! PLEASE add your comments to the open Github bug report! (https://github.com/openai/codex/issues/18778)) The approval system is utter pants and after 5m I'm already bored with multiple-choice options with ZERO visual difference. If I'm lucky I can see the "... in this project" and click that fucker immediately. Still work to do, but.. WORKING! (seriously! I just had to grant permission to write to a file IT CREATED around 2s ago! WTF! "User is frustrated with the constant confirmations, I will now attempt to work without causing further prompts". YES! GET THE FUCK ON WITH IT! Note: after me saying "Just get the fuck on with it!" .. ZERO permissions prompts! Even though that is explicitly a "No:" option. Never the less. works EVERY TIME! I have Gemini fixing up my HTML docs right now, for free! I always liked Gemini best for chat. The idea of it also being good at code is mighty appealing. Now it's updating the tooltips in my app! Editing my actual files, for free (this lasts about 5m, then my allowance runs out, well, shit!). BUT it's still dumb for code. I hit it with a fairly simple task and it still created two completely different "paths", when a) one was fine, with a simple param change, and b) the code already exists, and should have been re-used. But it's looking good. When the new models come out, I'M ON IT! Edited October 3 by corz TheDcoder 1 nothing is foolproof to the sufficiently talented fool..
corz Posted October 3 Author Posted October 3 Added when I started getting my AI to add "placeholder" comments above new functions, with a view to minimal future editing.. Quote ## Function Header Comments - Every comment header opens by stating its referent: what the function does, and if it returns anything non-trivial, exactly what that value is and what it means. Never describe a return value, a parameter, or a side effect without naming which one it is. - No private jargon: a first-time reader must understand every term from the block alone. Name the tool, the concept, or the format, or drop the phrase. - Never restate the signature in prose. Say what the signature cannot: purpose, contract, pitfalls. - Do *not* anthropomorphise any kind of function or data in your generated comments. Be precise: a value does not "sneak through", it "bypasses the check". A variable cannot "hold" data, it has no hands! It "contains" data. Similarly, files do not have "friends". Avoid *any* phrase which implies a physical manifestation where one does not exist. nothing is foolproof to the sufficiently talented fool..
corz Posted October 3 Author Posted October 3 Here's another thing AI is useful for, not specifically AGENTS.md-related, but you could definitely add something for this! No actual need though ... Sometimes I'll find really vague notes in my 2do files. I read it and think WTF did I mean there? I mean stuff like, "load from text file, one per line with @Date* token to fill the word. ???"<-- actual example. I feed this to Muse, saying, "What did I mean by that?" Muse scans entire codebase, does Sherlock Holmes number and returns with.. EXACTLY WHAT I MEANT! Not just the idea I was actually getting at, but all the actual code ready and waiting to implement and its tail is wagging (regardless of what I put in my AGENTS.md!!!), asking, "Switch me to build mode and I'll do it". I looked at that 2do entry like a dozen times before I thought, "Fuck it! AI can figure it out". Seconds later, mystery solved. That's wasted time, my AutoIt buddies! TheDcoder 1 nothing is foolproof to the sufficiently talented fool..
TheDcoder Posted October 3 Posted October 3 1 hour ago, corz said: OpenCode and Codex (which still lacks tabs for open sessions! I tried OpenCode today and I was severely disappointed by the lack of functionality in the UI, there's no way to edit permissions and it doesn't show me what it's doing behind the generic "Thinking" status text, I like being in control. I also didn't find an easy way to add MCPs etc., seems all of that is still done via manual json editing. @corz By the way, I tried to get Ternary Bonsai 2 (Qwen 3.8) to generate a cheatsheet for its own reference but it thought for a very long time (~30 minutes) and produced an obviously incomplete document by the end. My prompt was very simple: Quote Can you read the documentation and prepare a compact but dense cheatsheet to teach yourself for future SQLPage projects? You can save your work in a markdown file. Don't bother starting the server, just read the files directly. Any suggestions on what can be improved?
corz Posted October 3 Author Posted October 3 14 minutes ago, TheDcoder said: I tried OpenCode today and I was severely disappointed by the lack of functionality in the UI, there's no way to edit permissions and it doesn't show me what it's doing behind the generic "Thinking" status text, I like being in control. I also didn't find an easy way to add MCPs etc., seems all of that is still done via manual json editing. @corz By the way, I tried to get Ternary Bonsai 2 (Qwen 3.8) to generate a cheatsheet for its own reference but it thought for a very long time (~30 minutes) and produced an obviously incomplete document by the end. My prompt was very simple: Any suggestions on what can be improved? None of the AI interfaces are even close to perfect at this time. OpenCode, for me, is the best I've used and I see the frontier apps trying to copy it. (but still no tabs in Codex!) As for tips, as much as I'd love to dive in with "help", there are simply too many variables at play here. The bits I can see (your prompt) aren't "bad", per se, but could definitely be a *lot* better! I haven't played with the Bonsai, nor much with Qwen 3.8, but I *know* Qwen 3.6, and letting it loose with vague concepts like "the documentation", and "cheatsheet" would not fly, not if I was going away and leaving it to the task. I'm a huge believer in "steering", and this is my default <enter> action in OpenCode. I've seen entire includes explode and get rewritten after I added something like, "but not to the default model, of course". But if you plan to walk away, your prompt needs to be BULLETPROOF, as you are not there to steer. I tend to do a couple trial runs while I'm there, watch it think, see what trips it up, and then provide steering and examples where this tripping-up happens. "the documentation" better be in a format it understands, and in a location you have specified, and it has read access to the file, and so on, and that's just accessing the file! YES! I totally fucking hear you when it comes to adding shit. MCP servers? Get out the text editor baby! *sigh*. I console myself with the fact that it does the important stuff really well. I've worked with some cracking apps that can only be configured via text file, and indeed entire operating systems! Let's call it a work in progress. As to your "watching it think", use the CLI version with a local model and you can see every tiny detail! Even the stuff you wish couldn't! In the desktop version + frontier model you get vague thinking (and even that's encrypted and WILL NOT be decrypted, I've tried) but you can see what files it's working with and what commands it's running. I've stopped caring about the inner "thinking" too much, caring more about results. TheDcoder 1 nothing is foolproof to the sufficiently talented fool..
corz Posted Saturday at 10:40 PM Author Posted Saturday at 10:40 PM (edited) I've been messing with Codex a lot these last couple days. It is excellent at following the AGENTS.md (OMG! Did he just do that!?! 😁) I got an unexpected token reset (they are ALWAYS doing this! It's mindfuckery, but welcome) so fed it some work, lots of it. I made the mistake of clicking one of its code links and became humorously infuriated gazing at the lack of Syntax highlighting; funny, as I've done this before, and then thought, wait! AI apps be a building right now, they should be building with ME (and tangentially I suppose, YOU) in mind. I tried: Quote /suggest I wish your sidebar would syntax highlight AutoIt code! And it responded. So THAT didn't work! Then I asked it "how do I send feedback?" Quote It’s /feedback. Type it in the Codex composer and select the command from the menu, then send your AutoIt syntax-highlighting suggestion. I was wrong to treat /suggest as a feedback command. Official OpenAI documentation lists /feedback for product feedback; command availability can vary by app version. done! Thanks! Nice. Thank you for sending it! And NOTED! 😄 OH YOU MADE A BIG MISTAKE THERE BOYO! lol I’ll check the recent File mode label changes first and fix anything I can confirm. What did you spot? A brief description or screenshot will help me target the mistake. </me hits STOP> No! Breathe easy, the code is great. It was a joke, you letting me access this pipeline, as in, "I'll be sending ten thousand /feedback requests every day! Ha! I see it now. I walked straight into that one. 😄 YES YOU DID! lol Completely. 😂 NOT just a humorous story about my adventures in Codex but a CALL TO ARMS! In Codex: /feedback Please add AutoIt Syntax highlighting to your sidebar <insert your own thing> Let's get AutoIt right up the top of all future AI app combos! ;o) Edited Saturday at 10:42 PM by corz TheDcoder 1 nothing is foolproof to the sufficiently talented fool..
TheDcoder Posted Sunday at 12:42 AM Posted Sunday at 12:42 AM 1 hour ago, corz said: I've been messing with Codex a lot Sorry if you mentioned this before, but are you using it with the OpenAI subscription? I tried using it with my local llama.cpp server but apparently its OpenAI compatible API is too old.
corz Posted Sunday at 11:32 AM Author Posted Sunday at 11:32 AM 10 hours ago, TheDcoder said: Sorry if you mentioned this before, but are you using it with the OpenAI subscription? I tried using it with my local llama.cpp server but apparently its OpenAI compatible API is too old. I am using it with an OpenAI sub, yes. They gave me a free month "Plus" plan, which included three resets, so I got a lot done with it. I let it renew, thinking it was only $20/mo (around £15) but it turns out in the UK you pay £20/mo (darned fine print!), so compared to OpenCode's actual $10/mo it's looking less tasty. Sol (light) is a cheap to run and decent for most coding tasks, and it will run inside OpenCode too, but you need a Plus/Pro subscription to do it. OpenCode's Go package is way more bang for your bucks right now (you can run multiple models with multiple agents simultaneously inside OpenCode). I must say, having both at the same time has been super-productive, so long as you can keep up with managing their tasks. Really I need a third agent to deal with that! TheDcoder 1 nothing is foolproof to the sufficiently talented fool..
corz Posted Sunday at 04:42 PM Author Posted Sunday at 04:42 PM Another difference between Muse and Sol is that Sol will regularly underperform, compared to its "plan", whereas Muse will regularly overperform, compared to its plan. If I remember, I'll also explain why it's brilliant to have BOTH. This sounds highly subjective and it is, but Sol will tell you that it has the solution, and makes it look easy, but the final implementation is lacking, and it takes a bit (or a lot) of back-and-forth to make it *just* right. Muse, on the other hand, will layout every single detail of the plan (both models using identical AGENTS.md) and sometimes that's too much and you go fuck it! And get Sol to do it! But I've learned that this is usually a mistake. All those details up front enables you to *tweak* all those details, and with its massive context it will (grudgingly, it much prefers a stream of multiple choice questions. My response, "No, I'm not at school, let's discuss it") keep the whole plan in mind, along with all tweaks, and when you hit build mode, it will do it BUT, and here's where it overperforms and why choosing Sol over Muse is usually a mistake; during the task, it will often see underlying issues, tangential issues, false pathways and other stuff and just fix them as it goes. The end results are usually not just that one task done, but every related piece of code checked while it was at it, fixed and reported at the end; "I noticed that the caller's parameter's were.." or whatever, now fixed. It takes liberties, but only when it truly believes it will make the code better, more efficient. OpenAi models, on the other hand will just go ahead and implement shit because it thought it was a better way to do things, even if it runs totally opposite to your spec, alters user-facing controls, whatever. So back-and-forth afterwards, instead of Muse's back-and-forth before. Don't get me wrong, Sol will go into explicit details if pushed, and Muse, if pushed, will give you a simple summary and dive in, but these are their default behaviours, and if you work with one or the other a lot, model-specific AGENTS.md sections (like my Muse section) are not a bad idea. I'm considering a Sol section, to make it more Muse-Like. How do I know how to write an AGENTS.md section that some model will recognize and follow? I asked the model to do it! Muse proudly declared "I am Muse!" when asked and created my section, swearing that it would recognise that in the future, and it 100% does. Earlier today I wanted something done and Muse came up with a plan for me. I was dubious; I remembered reading some MSDN document a while back that contradicted his "fallback" proposal. I asked Sol the same question and its solution was cleaner, in line with what I had in mind. I got Sol to do it then told Muse to look at it. Initially it refused to believe that its "fallback" wasn't required, but when I pointed it to the Microsoft documentation page (provided by Sol) it relented and concurred (use experience to create AGENTS.md line!). Later I ask both for an implementation for a function that will return the number of lines in a file following specific rules. Sol, asked to to return with specific code, shows me its long-winded proposal. Muse, on the other hand, points out that I already have such a function in my own shared includes library and proposes a clean thin wrapper to use all its goodies as-they-already-are. This was the correct answer! Afterwards, Sol is very apologetic, "How could I have missed that!", but the fact remains, it missed it. Yesterday it built me an entire framework for capturing mouse hover over child controls. I respond, "Eh? This is native to Windows API" (Sol quickly rewrites code! *sigh* More token ...). All models can be dumb, no matter what you put in AGENTS.md. A single OpenCode Go sub gets you heaps of models, and you can use any other model to check your regular model's work. I'm using Sol/Astra with my Muse because I currently have access, but any two models with a mindful you in the middle gets you top-class code in the end. argumentum and TheDcoder 2 nothing is foolproof to the sufficiently talented fool..
corz Posted Sunday at 07:51 PM Author Posted Sunday at 07:51 PM 19 hours ago, TheDcoder said: Sorry if you mentioned this before, but are you using it with the OpenAI subscription? I tried using it with my local llama.cpp server but apparently its OpenAI compatible API is too old. This also sounds like you need to download and compile the latest sauce. nothing is foolproof to the sufficiently talented fool..
corz Posted Monday at 06:49 AM Author Posted Monday at 06:49 AM This is your last line, after you lay out your plan / goal / intentions. You need a clear prompt above this. Quote I believe in you, you got build mode, do your thing! The signal that you are handing over responsibility to the AI seems to kick in some deep programming, as does giving it immediate access to build mode, as does giving it the "freedom" to create. It will 100% follow whatever you put before that, but with extra *zest*, double-checking it's work, using extra error-checking. Switch on all output in OpenCode and you can watch the scripts it's creating to do work on the system, watch how they comprehensively improve when you tell it that you are relying on it to (or implying that it should) handle the task all by itself. This is for Muse, and probably other models, but especially Muse, which I'm getting to "know". I have found this to be the fastest route to producing quality code requiring the minimal amount of back-and-forth afterwards. It definitely employs more "thinking", and may take a few minutes extra, but Muse is basically free and this type of guidance, in my experience, is well worth it. Note: NOT For Kimi3. That fucker will burn through your 5h allowance in 5m doing any kind of investigation or double-checking. NEVER ask Kimi to look for bugs! TheDcoder and argumentum 1 1 nothing is foolproof to the sufficiently talented fool..
argumentum Posted Monday at 11:13 AM Posted Monday at 11:13 AM (edited) hmm, "Updates to Gemini models Flash and Pro will soon move to paid plans. You’ll still have unlimited access to the latest Flash-Lite, built for speed and intelligence." ..the financial struggle of "try AI, you'll love it". But it was the way to "..you'll need it". So if you liked the taste of Cocaine-AI(r)(tm)(🤪) ( soon to be Cocaine-SI ), the 1st was free to feel the rush, thereon after, you'll have to pay, and you'll pay 👹 Edited Monday at 01:44 PM by argumentum added video link TheDcoder and corz 1 1 Follow the link to my code contribution ( and other things too ). FAQ - Please Read Before Posting
corz Posted Monday at 04:32 PM Author Posted Monday at 04:32 PM (edited) One thing I've found myself doing a fair bit, is ask both models (yes, like Sith, always two there are!) to come up with a plan for the exact same thing. This is dynamite, and I should have mentioned it earlier. It's easy to see which model has a better handle on things; looks like Muse, because Sol will just give you a vague outline and expects a simple "yes" (why would a lowly human need details?). SO, you copy the plan Muse gave you (which contains every single detail of every proposed change) and hand this to Sol and say: Quote That's not a plan! THIS is a plan: <paste Muse plan verbatim> Sol then apologises and admits that its previous effort was vague and sketchy (but does not acknowledge the Crocodile Dundee reference). But then; and here's the good bit; it will evaluate your plan. It will be very keen to tell you about any errors and such as this was not its plan! So Sol will say something like, "I’d retain one detail for implementation: <exposes potential issue>", which you then feed back to Muse, who says, "That's a good idea, okay, final plan is now: <new plan>". NOW you have a plan, which any cheap model (Muse!) can carry out on-spec without issue. Edited Monday at 06:03 PM by corz argumentum 1 nothing is foolproof to the sufficiently talented fool..
corz Posted Monday at 05:57 PM Author Posted Monday at 05:57 PM If Muse is bamboozling you with its reply, here's a one-liner that will make a WORLD of difference (assuming you have a section something like mine in your AGENTS.md).. Quote Can I have that exact same information again, but with language that follows my AGENTS.md rules for Documentation Voice? WHAM! I was getting it to fix some of the "machine headers" that were sitting above "my" new functions. Before: Quote FunctionName() carried a second function's header ("Builds one fitted, resizable box window... Returns 0 on failure", true of nothing). Cut it; the true work-area paragraph stays, plus the real false path. After: Quote FunctionName() was wearing another function's header, the box-building paragraph with its "Returns 0 on failure" which was true of nothing. It is gone; the genuine work-area paragraph stays, now naming the real false path when no monitor answers. Okay, it's still not Walt Whitman but at least my human brain has a clue what it means. And of course, you can certainly improve on my shitty AGENTS.md, which would make the output better still. If you do that, paste it here, eh! nothing is foolproof to the sufficiently talented fool..
corz Posted Monday at 06:22 PM Author Posted Monday at 06:22 PM (edited) On 9/21/2026 at 8:36 AM, mLipok said: With ChatGPT codex this is taken automaticaly. I just noticed this! YES! This is something the agent/harness does. We do not need to manually send it! But the harness must attach this to EVERY prompt, along with whatever context you have already accumulated, which may be a lot. At some point, the model context simply drowns out the AGENTS.md, and its rules will no longer be "concrete", like they are in a fresh session. Those "rules" are way back at the start of the data the model receives, so old data gets less and less important as it always prioritises "recent" data, which to me seems like a flaw in the harness and indeed Codex (or maybe it's just Sol) seems less affected by this. Just how much actual data a specific model can handle before it starts to fall off that cliff depends on the model, but when it starts to happen, it's a thing you will notice, if you are watching for it. As I mentioned, Muse past 50% context is a mistake, and compacting, contrary to what's expected, just makes things worse. A fresh session, with perhaps a concise summary pulled over from the previous session, is always superior. Edited Monday at 06:24 PM by corz TheDcoder 1 nothing is foolproof to the sufficiently talented fool..
Recommended Posts
Create an account or sign in to comment
You need to be a member in order to leave a comment
Create an account
Sign up for a new account in our community. It's easy!
Register a new accountSign in
Already have an account? Sign in here.
Sign In Now