You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
- Updated Challenge-05 to use GitHub Spec Kit with real commands
- Added proper installation steps using uv tool install
- Included full workflow: constitution → spec → plan → tasks → implement
- Updated Coach guide with troubleshooting for uv installation
- Added reflection questions and learning resources
- Add Part 1: Build app with ad-hoc prompting (no structure)
- Add Part 2: Build same app with Spec Kit workflow
- Add Part 3: Compare results between approaches
- Remove installation instructions (moved to Challenge 0 prerequisites)
- Clarify that Spec Kit is one of many spec-driven approaches
- Update Coach guide to match new structure
The reason will be displayed to describe this comment to others. Learn more.
Pull request overview
Initial content drop for the new 076 – GitHub Copilot Cost Optimization What The Hack. It adds the student challenge materials, supporting resource assets (including a small NYC starter app), and coach guides/solutions (including a model-measurement harness).
Changes:
Added the full student challenge write-ups (Challenges 00–06) and the hack landing README.
Added student “Resources” content, including the CityScout NYC mini-app and Challenge 03 model-selection worksheets.
Added coach guides and solution assets (including measured results + a reproducible harness).
Reviewed changes
Copilot reviewed 50 out of 53 changed files in this pull request and generated 9 comments.
Agent Skills are discovered and loaded on demand when Copilot determines they are relevant; they do not require explicit invocation. This instruction teaches the wrong activation model and may invalidate the context-cost comparison. Describe the skill as relevance-triggered/on-demand instead. 076-GitHubCopilotCostOptimization/Student/Challenge-02.md:118
This path-specific instruction guidance contradicts the supported structure already described on line 28. Scoped instructions belong under .github/instructions/ and use applyTo frontmatter; placing /api/.github/ inside the source tree will not create the intended scope. 076-GitHubCopilotCostOptimization/Student/Challenge-03.md:34
This tells students to rerun anomalous results immediately after the first-shot rule forbids reruns. That changes the sample count and undermines a consistent comparison. Keep the measurement single-shot here; discuss repeated validation separately from the recorded run. 076-GitHubCopilotCostOptimization/Coach/Solution-02.md:98
Copilot does not scope repository instructions by walking nested .github/copilot-instructions.md files. The reference solution correctly uses .github/instructions/*.instructions.md with applyTo frontmatter, so this troubleshooting advice sends coaches toward an unsupported structure.
GitHub Copilot walks up the directory tree looking for `.github/copilot-instructions.md` files. Ensure:
- Directory structure is correct
- File is named exactly `.github/copilot-instructions.md` (not `.github/copilot-instructions-api.md` in the api folder)
All five challenge references are shifted: instruction scoping is Challenge 02, model selection is 03, session/tool hygiene is 04, compaction is 05, and spec-first development is 01. The incorrect mapping will make coaches point participants to the wrong prior material.
This issue also appears on line 40 of the same file.
This scoped rule requires location, but neither the student nor coach event fixtures contain that field, and the stated endpoint acceptance criteria require only name/date/price behavior. The completed solution therefore violates its own instructions; align the required fields with the provided schema.
The coach README says the harness defaults to n=4, and all published reference measurements use four trials, but the actual default is 50. Running the documented command without -n therefore incurs 12.5× as many API calls and can heavily rate-limit premium models.
Skills are selected automatically when their metadata matches the current task, not only when explicitly invoked; @workspace /new is also unrelated to loading the added Agent Skill. This coach explanation would reinforce the same incorrect activation model taught to students.
2. **GitHub Copilot Skills (conditional cost):**
- Only load when explicitly invoked
- Example: `@workspace /new` only loads creation templates when called
.devcontainer/devcontainer.json:8
This newly added default devcontainer is for hack 031 and opens /workspace/031-DevOpsWithGitHub, not the 076 materials introduced by this PR. Because it occupies the default .devcontainer/devcontainer.json location, users can launch an unrelated .NET 6 environment instead of the 076 TypeScript/Python configuration. Remove this file or replace it with the 076 configuration.
"name" : "DevOps with GitHub WTH in Codespaces",
// Or use a Dockerfile or Docker Compose file. More info: https://containers.dev/guide/dockerfile
"image": "mcr.microsoft.com/devcontainers/dotnet:0-6.0",
"workspaceFolder": "/workspace/031-DevOpsWithGitHub",
"workspaceMount": "source=${localWorkspaceFolder},target=/workspace,type=bind,consistency=cached",
076-GitHubCopilotCostOptimization/README.md:25
Challenge 01 explicitly excludes automated tests and frames the exercise as an observation, not proof that SDD reduces token burn. This summary instead promises a deterministic-control result the exercise neither implements nor requires. Align the overview with the actual direct-prompt-versus-spec comparison.
This guide does not use Spec Kit: it deliberately uses a single manually created sdd.md with no Spec Kit commands, constitution, plan, or tasks. Calling it “with Spec Kit” contradicts the student exercise and the rest of this coach guide.
# Challenge 01 - Spec-Driven Development with Spec Kit - Coach's Guide
These exact context-utilization thresholds are presented as deterministic behavior, yet the guide provides no measurement basis and later requires degradation after 70%. Quality degradation varies by model, task, and context composition, so a participant may run the experiment correctly and still fail the stated criterion. Label these as hypotheses/illustrative ranges and grade the recorded observation rather than a required outcome.
- 0-60% capacity: Quality remains high
- 60-70% capacity: First signs of degradation (inconsistencies)
- 70-85% capacity: Noticeable quality loss, forgotten requirements
- 85-100% capacity: Severe degradation, expensive recovery loops
These bullets contradict the explicit statement on lines 65–67 that the hack requires no Azure resources. Remove this leftover template content so coaches are not told to prepare nonexistent Azure dependencies.
This issue also appears on line 107 of the same file.
- Azure resources that will be consumed by a student implementing the hack's challenges
- Azure permissions required by a student to complete the hack's challenges.
Student/Resources/Challenge06 contains only data/sales.csv; the task specification and tests promised here are missing. Consequently, teams cannot know the acceptance criteria, run the required tests, or complete the capstone. Add token-golf-task.md and all referenced test/starter files to the resource folder. 076-GitHubCopilotCostOptimization/Coach/Solution-06.md:44
The same off-by-one challenge mapping is repeated in the required scorecard, so teams would attribute every optimization technique to the wrong challenge.
This resource-level configuration also opens a nonexistent Student/Resources/README.md, so no guide appears when the container starts. Point it to an existing resource README. .devcontainer/076-GitHubCopilotCostOptimization/devcontainer.json:23
The configured startup file does not exist: Student/Resources has no README.md. Codespaces will fail to open the intended guide. Point this to an existing resource entry page, such as the CityScout README.
The documented Token Golf task does not exist in Student/Resources/Challenge06, so the coach cannot perform this mandatory pre-work or establish par. Add the complete task/test assets before making this a critical prerequisite.
**You must complete the Token Golf task yourself before the hack:**
1. Implement the coding task specified in `/Student/Resources/Challenge06/token-golf-task.md`
2. Use optimization techniques from Challenges 1-5
This stray two-backtick marker is unmatched and renders as unexplained literal content in the instruction file. Remove it. 076-GitHubCopilotCostOptimization/Coach/README.md:109
A complete “Suggested Hack Agenda” and “Repository Contents” already appear above; this second section is untouched template guidance with sample days and duplicated headings. Remove the duplicate template block (lines 107–131) before publishing the coach guide.
## Suggested Hack Agenda (Optional)
_This section is optional. You may wish to provide an estimate of how long each challenge should take for an average squad of students to complete and/or a proposal of how many challenges a coach should structure each session for a multi-session hack event. For example:_
These language-suffixed filenames and nested .github/copilot-instructions.md paths do not provide automatic language/path scoping in the documented VS Code workflow. Moving rules there can stop them being applied rather than reduce their scope. Use .github/instructions/*.instructions.md with YAML applyTo globs, as the supplied solution's api.instructions.md already does. Correct the troubleshooting advice at lines 96–98, the reference tree at lines 140–142, and the same tip in Student/Challenge-02.md:118 as well.
Align event validation with the available schema fields
The supplied events have name, date, and price, but no location. This reference instruction requires a missing field, while the accompanying skill says to exclude incomplete entries. Following both could discard valid fixture records or prompt unnecessary schema changes. Align the rule with the fields used by the endpoint.
Replace word-order scoring with constrained answer validation
Word order does not identify the recommendation: this checker rejects “You shouldn't walk; drive the car to the wash” and accepts “Driving is unnecessary; walk instead.” Both cases reproduce, so ordinary explanatory wording can invert the reported pass/fail results. Preserve raw answers for manual assessment, or use a constrained answer format and validate its label rather than scoring unrestricted prose by substring position.
Align model selection with current usage-based billing rates
Challenge 00 requires usage-based billing, but these model definitions use included requests and premium-request multipliers. Under the linked organization/enterprise usage-based billing documentation, chat is charged in AI credits using model token rates, so those multipliers are not the cost metric for this exercise. Select models by their current input, cached-input, and output rates, and compare measured credits. Update the same tier-selection guidance in the resource README and Coach/Solution-03.md, including retry-cost calculations.
Restore initial files and use a fresh Run B conversation
Run A has already implemented all four features, but Run B does not explicitly restore the starting files or start an independent conversation. Reusing the completed code or chat history changes the task and makes the credit comparison unreliable. Require the same initial code and configuration, plus a fresh conversation, before Run B.
Align pickleball planning materials with the grading rubric
Student/Challenge-03.md Part 2 assigns a pickleball implementation plan, scored for clarifying questions, risk depth, specificity, and trustworthiness. These packaged prompts still assign the car-wash question; measurements-template.md records walk/drive correctness, and Coach/Solution-03.md validates the car-wash results. Students and coaches would perform and assess different experiments. Align the resource README, prompts, worksheet, and coach guidance with the planning rubric, and clearly label the car-wash harness/results as a separate optional experiment if retained.
Visible response length cannot establish total token use or credit cost: input, cached-input, and hidden reasoning tokens are not represented by that text. This fallback therefore cannot validate the worksheet's token ratios or retry-cost calculations. Mark unavailable usage as unmeasured, keep verbosity as a separate qualitative observation, and label harness/playground data as a separate reference demonstration. Apply the same correction to Student/Challenge-03.md:30, the measurement template's introduction, and this guide's later token-visibility advice.
The supplied events contain name, date, and price, but no location. This instruction therefore marks every fixture incomplete; combined with the skill's instruction to exclude incomplete entries, it can lead students to filter out all events. Match the reference instruction to the fixture schema.
Creating this skill does not make it explicit-invocation-only: Copilot can select it automatically when a task matches its description. The supplied free-events skill specifically matches the baseline task. Explain the cost distinction as discoverable name/description metadata versus a conditionally loaded skill body. Update the repeated tip at line 115 and Coach/Solution-02.md:19-20 too; @workspace /new is not an example of a SKILL.md skill. See GitHub's skill documentation.
Model selection uses obsolete premium-request multipliers
These model-selection instructions use legacy premium-request multipliers, although Challenge 00 requires usage-based billing. Under that billing system, model-specific input, output, and cached-token rates determine credit cost; a no-multiplier included tier is not the relevant distinction. Select the comparison models using current token prices and calculate costs from those rates and measured usage. Align the resource README, coach terminology, and multiplier-based retry guidance too. See GitHub's billing transition documentation.
This permits rerunning an odd result immediately after the first-shot-only rule. Selective reruns produce inconsistent sample sizes and can replace surprising observations with preferred results. Keep the first-shot measurement, and record any additional trials separately using the same trial count for every model/task pair.
Completion criteria require outcomes the experiment cannot guarantee
These completion criteria require Run A to show failures and Run B to cost less. Neither result is guaranteed: model responses vary, and compaction itself consumes tokens. A student can conduct the experiment correctly without observing either result. Assess whether students document and explain the measured outcomes rather than requiring a predetermined winner.
Updated instructions for running model selection experiments with GitHub Copilot Chat, including steps for recording answers and token usage.
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
The reason will be displayed to describe this comment to others. Learn more.
change to "Implement the complete application on the local machine."
This branch has not been deployed
No deployments
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This is the initial pull request for the new GitHub Copilot Cost Optimization WTH