npm package discovery and stats viewer.

Discover Tips

  • General search

    [free text search, go nuts!]

  • Package details

    pkg:[package-name]

  • User packages

    @[username]

Sponsor

Optimize Toolset

I’ve always been into building performant and accessible sites, but lately I’ve been taking it extremely seriously. So much so that I’ve been building a tool to help me optimize and monitor the sites that I build to make sure that I’m making an attempt to offer the best experience to those who visit them. If you’re into performant, accessible and SEO friendly sites, you might like it too! You can check it out at Optimize Toolset.

About

Hi, 👋, I’m Ryan Hefner  and I built this site for me, and you! The goal of this site was to provide an easy way for me to check the stats on my npm packages, both for prioritizing issues and updates, and to give me a little kick in the pants to keep up on stuff.

As I was building it, I realized that I was actually using the tool to build the tool, and figured I might as well put this out there and hopefully others will find it to be a fast and useful way to search and browse npm packages as I have.

If you’re interested in other things I’m working on, follow me on Twitter or check out the open source projects I’ve been publishing on GitHub.

I am also working on a Twitter bot for this site to tweet the most popular, newest, random packages from npm. Please follow that account now and it will start sending out packages soon–ish.

Open Software & Tools

This site wouldn’t be possible without the immense generosity and tireless efforts from the people who make contributions to the world and share their work via open source initiatives. Thank you 🙏

© 2026 – Pkg Stats / Ryan Hefner

@learning-commons/evaluators

v1.1.1

Published

TypeScript SDK for Learning Commons educational evaluators

Readme

@learning-commons/evaluators

npm version

TypeScript SDK for Learning Commons evaluators — sixteen LLM-backed evaluators for the complexity of text students read, the quality of feedback they receive, and the alignment of math items to standards.

Requires Node 20.19+ or 22.12+ (^20.19.0 || >=22.12.0) — the CommonJS build needs require(esm), which Node 21.x and 22.0-22.11 lack.

The SDK is ESM-first and the examples below use top-level await, so a TypeScript project needs "type": "module" in package.json, @types/node, and:

{
  "compilerOptions": {
    "module": "nodenext",
    "moduleResolution": "nodenext",
    "target": "es2022",
    "strict": true,
    "types": ["node"]
  }
}

You will also want the type packages:

npm install -D typescript @types/node @types/json-schema

@types/json-schema is needed because ai's own dependency chain imports json-schema untyped; without it tsc reports TS7016 from inside @ai-sdk/provider, not from this package. This SDK's declarations typecheck cleanly on their own, so skipLibCheck is not required on their account — but adding it is the other way to silence that peer's gap.

To run a snippet, use a TypeScript runner. The config above is noEmit-shaped, so there is no build step to run:

npx tsx quickstart.ts

A runner is the portable choice. Recent Node versions strip types themselves, so node quickstart.ts works there too, but the 20.19 floor above cannot.

CommonJS consumers can require the package but will need to rewrite the top-level await in these snippets. Two caveats there: the CommonJS build requires ESM-only packages, so it needs require(esm) (hence the Node range above), and Jest's default CommonJS transform cannot load it at all: SyntaxError: Cannot use import statement outside a module, from inside ai. Add those packages to transformIgnorePatterns, or run Jest in ESM mode:

transformIgnorePatterns: ['node_modules/(?!(ai|p-limit|syllable|text-readability)/)'],

Each evaluator's argument and payload types are exported — VocabularyComplexityInput, VocabularyComplexityResult and so on — so you can name them in your own signatures. Both are generated from the evaluator's contract, so an input's declared values are a literal union rather than string:

const ok: PurposeClarityInput = { text, grade_level: "5" };
const typo: PurposeClarityInput = { text, grade_level: "fifth" };
//                                       ^ not assignable to '"3" | "4" | … | "12"'

A grade arriving as a string — from a form, a database, a request body — needs narrowing before it fits. Validate it where the external value enters your system, which is where the check belongs; the evaluator still rejects a bad value at run time either way.

Installation

Install the SDK alongside Vercel AI and zod:

npm install @learning-commons/evaluators ai zod

ai and zod are peer dependencies, so your project owns their versions. zod must be version 4 — the evaluator schemas this package exports are zod 4 values. Installing against zod 3 fails at install time, naming zod:

npm error ERESOLVE unable to resolve dependency tree
npm error Found: [email protected]

Forcing past it with --legacy-peer-deps or --force installs zod 3 anyway, and the build then fails on our declarations instead — Namespace '…/zod/v3/external' has no exported member 'core'.

If your project already depends on zod 3, widen your own range first. The install above will not do it for you, because npm install zod honours the ^3 you already declared:

npm install zod@^4          # then re-run the install above

That is a major upgrade of a dependency you own, so review zod's own 3 to 4 notes for your call sites. Nothing needs fixing at this SDK's call sites; the schemas it exports are zod 4 values, and once there is a single zod 4 copy they compose with yours directly.

The floor is zod@^4.1.8 rather than ^4.0.0 because ai@7 itself requires zod@^3.25.76 || ^4.1.8.

Next, install the provider adapter(s) for the evaluators you plan to run — the table below gives each evaluator's provider:

npm install @ai-sdk/openai     # OpenAI
npm install @ai-sdk/google     # Google Gemini
npm install @ai-sdk/anthropic  # Anthropic

Quickstart

import { GradeLevelAppropriatenessEvaluator } from "@learning-commons/evaluators";

const evaluator = new GradeLevelAppropriatenessEvaluator({
  googleApiKey: process.env.GOOGLE_API_KEY,
});

const { result, metadata } = await evaluator.evaluate({
  text: "The cat's out of the bag now.",
});

console.log(result.grade_band); // a CCSS band, e.g. "2-3"
console.log(result.alternative_grade_band); // the band reachable with scaffolding
console.log(result.scaffolding_needed); // what that band would need
console.log(metadata.model); // "google:gemini-3.6-flash"

Every evaluator resolves to the same three-part envelope, so generic code works across all sixteen:

{
  evaluator: string; // registry id, e.g. "text_complexity.ela_reading.vocabulary_complexity"
  result: TResult; // the evaluator's own payload, exactly as its output schema declares it
  metadata: {
    model: string; // "provider:model" that ran; "a+b" when several did
    processingTimeMs: number;
    tokenUsage: {
      inputTokens: number;
      outputTokens: number;
    }
  }
}

result is the model's structured output with keys and values unaltered, so the payload is identical across our SDKs. Payloads carry more than the verdict, and how much varies by evaluator: Background Knowledge Demands returns identified_topics, curriculum_check, assumptions_and_scaffolding and friction_analysis; Meaning Directness returns conventionality_features, grade_context and instructional_insights; Organizational Structure, Purpose Clarity and Reference Knowledge Demands each return a nested details; Vocabulary Complexity returns tier_2_words, tier_3_words, archaic_words and other_complex_words as comma-joined strings. Only Sentence Structure returns just the score and reasoning. Each evaluator exports its payload type (VocabularyComplexityResult and so on) — that is the authoritative shape.

The feedback family's quality_score is the number 0 or 1, and its key_features gives each criterion a { met: 0 | 1, justification: string }, where met is a number, not a boolean. key_features is a fixed struct, not an index-signature map, so key_features[name] for a string name is a type error; and the criterion keys differ per evaluator, e.g.

RevisionAccuracy     accurate_task_assessment, revision_need_identification,
                     appropriate_signal_when_task_complete
ToneAppropriateness  neutral_professional_language, targets_work_not_student,
                     praise_proportionate_to_work

To enumerate them for any evaluator, read its exported schema: Object.keys(ToneAppropriatenessOutputSchema.shape.key_features.shape).

metadata.model names every model that ran, joined by + when an evaluator uses more than one — so a multi-step evaluator reports e.g. openai:gpt-4o-…+openai:gpt-4.1-…. Which models run can also depend on the input: Vocabulary Complexity takes a different branch for grades 3-4 than for 5-12. The Default provider column below is what each evaluator's contract declares, and you must supply a key for every provider listed — not a promise about which one serves a given call: construction validates the union of keys an evaluator could need across all its branches, so Vocabulary Complexity demands both keys even at a grade where only one provider runs. Model strings come from each evaluator's contract and change with it, so treat metadata.model as the record of what actually ran rather than something to assert on. When you need one comparable value per evaluation regardless of evaluator, use readOutcome:

import {
  readOutcome,
  VocabularyComplexityEvaluator,
} from "@learning-commons/evaluators";

const evaluation = await new VocabularyComplexityEvaluator({
  googleApiKey: process.env.GOOGLE_API_KEY,
  openaiApiKey: process.env.OPENAI_API_KEY,
}).evaluate({ text, grade_level: "5" });

const { score, reasoning } = readOutcome(
  evaluation,
  VocabularyComplexityEvaluator.metadata.outcome,
);

score is always a string, or undefined when the evaluator declares no single verdict. That matters for the feedback family, whose verdict is the number 0 or 1: readOutcome returns "0", and "0" is truthy in JavaScript. Compare explicitly — score === "1" — or read the payload field directly (evaluation.result.quality_score, a real number) when you want arithmetic. The second argument's type is exported as DeclaredOutcome, so you can write one helper across several evaluators.

Evaluators

Text complexity — how demanding a text is for a given grade. Each takes { text, grade_level } and returns a complexity_score on a four-level scale — slightly_complex, moderately_complex, very_complex, exceedingly_complex — with reasoning. Purpose Clarity's complexity_score has a fifth possible value, more_context_needed, so a switch over the four above will not be exhaustive for it.

| Evaluator | Grades | Default provider | Docs | | ------------------------------------- | ------ | ---------------- | ----------------------------------------------------------------------------------------------------------- | | BackgroundKnowledgeDemandsEvaluator | 3–12 | Google | Link | | MeaningDirectnessEvaluator | 3–12 | Google | Link | | OrganizationalStructureEvaluator | 3–12 | Google | Link | | PurposeClarityEvaluator | 3–12 | Google | Link | | ReferenceKnowledgeDemandsEvaluator | 3–12 | Google | Link | | SentenceStructureEvaluator | 3–12 | OpenAI | Link | | VocabularyComplexityEvaluator | 3–12 | Google + OpenAI | Link |

Grade band — takes { text } only, and determines the grade rather than judging against one. Returns grade_band, alternative_grade_band, scaffolding_needed, reasoning. Bands are K-1, 2-3, 4-5, 6-8, 9-10, 11-12 — spans on the CCSS text-complexity scale, not single grades.

| Evaluator | Grades | Default provider | Docs | | ------------------------------------ | ------ | ---------------- | ---------------------------------------------------------------------------------------------------------- | | GradeLevelAppropriatenessEvaluator | K–12 | Google | Link |

Feedback quality — judges a teacher comment on a student's writing. Each takes { student_text, feedback_text } and returns a binary quality_score with reasoning, key_features and proposed_adjustment.

| Evaluator | Grades | Default provider | Docs | | ------------------------------------- | ------ | ---------------- | ---------------------------------------------------------------------------------------------------- | | RevisionAccuracyEvaluator | 6–12 | OpenAI | Link | | RevisionActionabilityEvaluator | 6–12 | OpenAI | Link | | RevisionManageabilityEvaluator | 6–12 | OpenAI | Link | | StrengthAcknowledgmentEvaluator | 6–12 | OpenAI | Link | | StudentResponseSpecificityEvaluator | 6–12 | OpenAI | Link | | ToneAppropriatenessEvaluator | 6–12 | OpenAI | Link | | WithholdingAnswersEvaluator | 6–12 | OpenAI | Link |

Standards alignment — checks a math item against a standard, component by component.

| Evaluator | Grades | Default provider | Also needs | Docs | | --------------------------------- | ------ | ---------------- | ----------------------------------------- | ---------------------------------------------------------------------------------------------------------- | | MathStandardsAlignmentEvaluator | K–12 | Anthropic | learningCommonsApiKey (Knowledge Graph) | Link |

import {
  MathStandardsAlignmentEvaluator,
  Jurisdiction,
} from "@learning-commons/evaluators";

const { result } = await new MathStandardsAlignmentEvaluator({
  anthropicApiKey: process.env.ANTHROPIC_API_KEY,
  learningCommonsApiKey: process.env.LEARNING_COMMONS_API_KEY,
}).evaluate({
  question: "A playground is shaped like an L. What is its area?",
  statement_code: "3.MD.C.7.d",
  jurisdiction: Jurisdiction.MultiState,
});

console.log(
  `${result.aligned_count}/${result.total_count} learning components aligned`,
);

result carries one verdict per learning component the standard declares, so you can see which part of the standard an item does and does not reach:

{
  "statement_code": "3.MD.C.7.d",
  "learning_components": [
    {
      "identifier": "29c42ed3-da20-5288-a4b1-c72989c97fe4",
      "description": "Find the area of rectilinear figures by decomposing them into non-overlapping parts and finding the area of each part",
      "reasoning": "The question presents an L-shaped playground explicitly described as two rectangular parts with no overlap...",
      "aligned": true,
      "feedback": "Students must decompose the L-shape into two rectangles, find each area, and sum them.",
    },
    // ...one entry per learning component
  ],
  "aligned_count": 3,
  "total_count": 3, // a total_count of 0 means the standard has no components authored yet
}

Each evaluator is also available as a function — evaluateGradeLevelAppropriateness(input, config) and so on — for callers who would rather not hold an instance.

Discovering evaluators

Every evaluator is listed in a registry, keyed by the registry id that appears on each result:

import { getEvaluators, getEvaluator } from "@learning-commons/evaluators";

for (const { id, name, supportedGrades } of getEvaluators()) {
  console.log(`${name} (${id}) — grades ${supportedGrades.join(", ")}`);
}

// Renamed ids still resolve, so a stored result stays identifiable.
getEvaluator("conventionality")?.name; // "Meaning Directness Evaluator"

Both return metadata — id, stableId, idHistory, name, description, supportedGrades, defaultProviders, requiredCredentials, and outcome where the evaluator declares a single verdict. requiredCredentials lists only non-LLM services — it is ["learning_commons_api_key"] for math standards alignment and [] for the other fifteen, so it is not the answer to "which keys does this need". Provider keys follow defaultProviders: ["google"] means supply googleApiKey. To run an evaluator, import it by name: the metadata does not tell you which named inputs it takes, and each evaluator's are different.

Configuration

Every evaluator takes the same options:

| Option | Purpose | | --------------------------------------------------- | -------------------------------------------------------------------------------------------------------------- | | googleApiKey / openaiApiKey / anthropicApiKey | Keys for the providers the evaluator uses | | learningCommonsApiKey | Authorizes Learning Commons API calls, such as the Knowledge Graph | | modelOverride | Run every call on a different { provider, model }; provider is the exported Provider enum | | llmProvider | Bring your own provider (see below) | | maxRetries | Retries per failed call (default 2, so 3 attempts) | | telemetry | true, false, or TelemetryOptions (default on) | | logger / logLevel | Inject a logger, or set the console logger's level with the exported LogLevel enum (default LogLevel.WARN) |

TelemetryOptions and Logger are both exported. Their shapes:

interface TelemetryOptions {
  enabled?: boolean; // default true
  learningCommonsApiKey?: string; // set: events are attributed to you. unset: anonymous
}

interface Logger {
  // LogContext is { evaluator?, operation?, error?, ...unknown }
  debug(message: string, context?: LogContext): void;
  info(message: string, context?: LogContext): void;
  warn(message: string, context?: LogContext): void;
  error(message: string, context?: LogContext): void;
}

Telemetry never includes the text you evaluate. The event carries its length, the evaluator, the grade level you passed, the provider, latency, status, any error class name, token counts, and the SDK version.

Telemetry failures are logged at warn and never affect an evaluation, so a restricted-egress environment will see a warning per call at the default level; telemetry: false silences it.

Provider and LogLevel are enums, not strings — provider: "google" and logLevel: "ERROR" do not compile. The members are Provider.OpenAI, Provider.Google, Provider.Anthropic, and LogLevel.DEBUG, INFO, WARN, ERROR, SILENT:

import {
  Provider,
  LogLevel,
  VocabularyComplexityEvaluator,
} from "@learning-commons/evaluators";

new VocabularyComplexityEvaluator({
  googleApiKey,
  openaiApiKey,
  modelOverride: { provider: Provider.Google, model: "gemini-2.5-flash" },
  logLevel: LogLevel.ERROR,
});

Keys are read from this object only. There is no environment-variable fallback: setting GOOGLE_API_KEY in the environment does not satisfy googleApiKey, and omitting a key an evaluator needs throws ConfigurationError at construction. (The evaluators-batch command is the exception — it does read the environment, as its --help describes.)

Bring your own provider

Pass any object implementing LLMProvider and the evaluator routes every call through it, skipping the built-in API-key adapters entirely — useful for Google Vertex AI, Amazon Bedrock, an AI-SDK gateway, or an eval framework's model system. No provider API keys are needed.

Three members are required, and the two methods are not symmetrical: generateStructured takes one object and must return a model, while generateText takes positional arguments and must not. messages includes the system turn, which most backends want as a separate argument. A temperature of null means send no temperature at all — some models reject an explicit value — so forward it conditionally rather than coercing it with ?? undefined.

import { generateText, Output } from "ai";
import { openai } from "@ai-sdk/openai";
import {
  SentenceStructureEvaluator,
  type LLMProvider,
  type Message,
} from "@learning-commons/evaluators";

/** `ai` takes the system turn as its own option; leaving it in `messages` is rejected. */
function split(messages: Message[]) {
  const system = messages.find((m) => m.role === "system");
  return {
    ...(system ? { system: system.content } : {}),
    messages: messages.filter((m) => m.role !== "system"),
  };
}

const myProvider: LLMProvider = {
  // "provider:model" — this is what the SDK reports as metadata.model.
  label: "byo:gpt-4o-mini",

  async generateStructured({ messages, schema, temperature }) {
    const started = Date.now();
    const { output, usage } = await generateText({
      model: openai("gpt-4o-mini"),
      ...split(messages),
      output: Output.object({ schema }),
      ...(temperature != null ? { temperature } : {}),
    });

    return {
      data: output,
      model: "gpt-4o-mini",
      usage: {
        inputTokens: usage.inputTokens ?? 0,
        outputTokens: usage.outputTokens ?? 0,
      },
      latencyMs: Date.now() - started,
    };
  },

  async generateText(messages, temperature) {
    const started = Date.now();
    const { text, usage } = await generateText({
      model: openai("gpt-4o-mini"),
      ...split(messages),
      ...(temperature != null ? { temperature } : {}),
    });

    return {
      text,
      usage: {
        inputTokens: usage.inputTokens ?? 0,
        outputTokens: usage.outputTokens ?? 0,
      },
      latencyMs: Date.now() - started,
    };
  },
};

const evaluation = await new SentenceStructureEvaluator({
  llmProvider: myProvider,
}).evaluate({
  text: "The dog ran. It was fast. The children laughed at the sight of it.",
  grade_level: "5",
});
// evaluation.metadata.model === "byo:gpt-4o-mini"

schema is a Zod schema, and because zod is a peer dependency it is an instance of your zod: the same copy your own code imports, which is what makes Output.object({ schema }) above compile without conversion. A backend that needs another form can convert it with zod's own helpers (z.toJSONSchema(schema)). An object missing any of the three members throws ConfigurationError.

It is mutually exclusive with modelOverride: setting both throws ConfigurationError.

Batch

@learning-commons/evaluators/batch runs a CSV of rows through a family of evaluators, with concurrency and retries. It returns results; it writes nothing. Use renderOutputs to format them, or the command below, which writes the files for you.

import { BatchEvaluator, parseCSV } from "@learning-commons/evaluators/batch";

const rows = parseCSV("./input.csv"); // a path, not CSV text

const output = await new BatchEvaluator({
  googleApiKey: process.env.GOOGLE_API_KEY,
  openaiApiKey: process.env.OPENAI_API_KEY,
  concurrency: 3,
}).evaluate(rows, "text-complexity", {
  onProgress: (result) => console.log(result.evaluatorId, result.status),
});

console.log(
  `${output.summary.successful}/${output.summary.totalTasks} succeeded`,
);

The same engine is installed as the evaluators-batch command, which writes the files for you:

npx evaluators-batch input.csv --family text-complexity --output-dir ./results -y

See the batch evaluator README for all evaluator families and their CSV columns, the full flag list, and the rest of the programmatic API (e.g., getFamily, renderOutputs, and the formatters).

Errors

Errors are grouped by fault domain, so you can catch by who is at fault rather than by individual failure. All extend EvaluatorError.

| Class | Meaning | | ---------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | ConfigurationError | The SDK was set up wrong — missing key, conflicting options, a model the provider rejects | | InputValidationError | The input was rejected before any model ran; StandardNotFoundError is a subclass | | EvaluationError | The evaluation ran but could not be completed; LLMOutputProcessingError is a subclass | | DependencyError | Something the SDK depends on failed. Subclasses: AuthenticationError, RateLimitError, NetworkError, RequestTimeoutError, LLMProviderError, KnowledgeGraphError |

A failed evaluation is also logged before it is thrown, so a caught error still prints an [ERROR] block at the default level. Pass logLevel: LogLevel.SILENT if you would rather report failures yourself.

try {
  await evaluator.evaluate({ text, grade_level: "5" });
} catch (error) {
  if (error instanceof RateLimitError) retryLater();
  else if (error instanceof DependencyError) reportUpstreamOutage(error);
  else if (error instanceof InputValidationError) fixTheRow(error);
  else throw error;
}

Documentation

Full reference at our docs site. Upgrading from an earlier major version? See MIGRATION.md.

License

MIT