Tag: big-tech

  • Values in AI

    Daniel Schmachtenberger has made the argument:

    1. All technologies embody value systems
    2. Some technologies are obligate in a competitive environment

    The example of his that stuck with me was the plough: many cultures were animistic (a belief in the spirit of the animal), but after the scaling up of agriculture enabled by the plough, most weren’t. The plough’s enablement of large-scale agriculture likely shifted societies toward sedentism (vs nomadism) and surplus, altering spiritual relationships with animals as they became tools for labor. The perspective shift — the value it encodes — is embedded in the technology.

    The plough is also obligate. If one group uses it and other doesn’t, the group that does will be able to farm more per person. That surplus enables for more specialization, which yields an advantage either in terms of trading or conflict. If the second group doesn’t adopt the plough they will be taken-over, outgrown, or both, by the first.

    AI may well be an obligate technology, which forces us to make deliberate ethical choices about its deployment and values. We are in the early stages of seeing that with software development. That’s going to change the nature of certain careers: changing what the day-to-day work looks like and impacting demand for software engineers. That isn’t necessarily negative: it will depend on the opportunities that replace the current ones. It also isn’t neutral: our approach to AI, how we deploy it, how it is used are all a series of choices that embed values.

    Some of those values are encoded into the models by the training data and loss functions, some are encoded in the systems engineering, the choices of which tasks to apply it to, which interactions to explore and so on, and some are explicitly engineered in through fine tuning and reinforcement learning.

    One way of looking at those values is through the study of ethics, how to live in a just way. This is a core topic for philosophers. One example is Kant’s Categorical Imperative, which requires actions to follow maxims that could be universally applied without contradiction, ensuring rational consistency.

    It’s somewhat akin to asking the question: Would I still support this if I knew everyone else would act this way? Further, would I support this action if knew I would be born again randomly into the world, maybe in a much different situation than my one now?

    The proliferation of useful AI agents adds a somewhat realistic flavor to the question: if, in the future, you are dependent on systems constrained by these specific guidelines or rules , are you happy about that?

    Kantian (or deontological) thinking is far from the only ethical system. A lot of thinking about AI ethics has been consequentialist. Consequentialism is practical: the “goodness” of an action is whether it results in a good outcome! Inherently we judge AI training (at least for RL and supervised learning) by the achievement of the outcome encoded in a loss function, reward function or similar. Stuart Russel (of & Norvig fame from university courses of my youth) has written about “provably beneficial” AI where the AI maximizes a human-involved reward signal (a little like the Assistant Games pattern we discussed before).

    The downside of all this is well documented — Nick Bostrom’s famous paperclip maximizer thought experiment is an AI that achieves the objective, but in a way that was undesirable. A more benign but annoying example might be a cleaning robot that pushes everything outside the house in order to make it tidier. Because outcome-based rules just judge the what, and now the how, they can also encourage power-seeking (as called out by Bostrom) in order to better achieve objectives.

    standard forms of consequentialism recommend taking unsafe actions when such acts maximize expected utility. Adding features like risk-aversion and future discounting may mitigate some of these safety issues, but it’s not clear they solve them entirely.

    Deontology and safe artificial intelligence – William D’Allesandro

    Anthropic’s constitutional AI approach can be seen as a blend of approaches; the constitution is a set of principles that can be used by another AI to criticize and improve output in response to requests:

    As AI systems become more capable, we would like to enlist their help to supervise other AIs. We experiment with methods for training a harmless AI assistant through self-improvement, without any human labels identifying harmful outputs. The only human oversight is provided through a list of rules or principles, and so we refer to the method as ‘Constitutional AI’. 

    The training still ultimately uses a form of reinforcement learning (which is inherently consequentialist), but the reward is given according to how well the outputs adhere to the constitutional principles.

    A more recent philosopher, Derek Parfitt, argued that all moral systems were hill climbing towards a shared perspective, and you can evaluate an action on multiple in order to gain confidence. For example, when considering an option, you could ask:

    a) Would it maximize overall good? (consequentialist)
    b) Could everyone rationally will it? (Kantian)
    c) Could anyone reasonably reject it? (contractualist1)

    “Rationally” here is doing a bit of work: it means “with reasoning”, as in there is a chain of thought that can support and justify the decision.

    Part of the challenge with rationalism is that part of the reward signal here is coming from human raters. We have seen this play out with LMSys where models which are “friendlier” score better, and in a more extreme version in the ChatGPT 4O misalignment where the model became excessively sycophantic in a way that resulted in better rewards in short doses, and didn’t impact any of the quantitative evaluations, despite being an overall negative to the experience.

    As we move into more agentic systems we often have fewer tools to evaluate or make visible the values we are encoding, but we are still doing it!

    For example. Google’s recent AlphaEvolve project uses Gemini underneath, which is an LLM that can be evaluated and aligned. But on top of that it uses an evolutionary search scheme (another reminder of Rich Sutton’s bitter lesson) to generate different prompts and evaluations and iterate towards a new, externally defined goal: in that case generating better algorithms and code. We are searching for superior outcomes, but that search itself is -somewhat unconstrained by other values: it’s a more consequentialist approach.

    The current crop of agentic coding tools often recommends encoding preference data into a project specific file. For example, Claude Code recommends a CLAUDE.md file

    • Include frequently used commands (build, test, lint) to avoid repeated searches
    • Document code style preferences and naming conventions
    • Add important architectural patterns specific to your project
    • CLAUDE.md memories can be used for both instructions shared with your team and for your individual preferences.

    While it presents them as memory, the idea here is to guide the choices of the model in a way that aligns with the principles by which the project being modified is managed.

    OpenAI have published work in allowing hierarchies of instruction: The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions | OpenAI

    we argue that one of the primary vulnerabilities underlying these attacks is that LLMs often consider system prompts (e.g., text from an application developer) to be the same priority as text from untrusted users and third parties. To address this, we propose an instruction hierarchy that explicitly defines how models should behave when instructions of different priorities conflict. 

    As well as using a single model that can incorporate different safeguards, we can use models themself to verify actions and outputs. Verification is generally an easier problem than generation, so a model that is unable to consistently follow a set of principles may still be able to validate whether a given example does or does not follow them.

    LlamaGuard is a good example of this kind of system, built and released by Meta’s GenAI team alongside Llama. One example of seeing this process in the wild is OpenAI’s safety systems on 4O image generation. Inherently agentic, 4O can generate image ideas, then the image itself. Despite the model having constraints on it, it will happily generate things which violate OpenAI’s content policy, necessitating a monitoring model that whisks them away before a use can access a violating image.

    If AI becomes an obligate technology, we will benefit from encoding values intentionally, balancing outcomes, universal principles, and fairness. The challenge is ensuring these choices reflect the world we want, not just the one we’re competing in.

    1. Another theory of ethics that weights mutuality heavily: it’s frames ethical considerations as something derived between people rather than just based on outcomes or on abstract principles. Its featured particularly in Scanlon’s What We Owe to Each Other for those, like me, who get all of their ethical understanding from watching The Good Place ↩︎

  • LSP & Standards

    https://www.michaelpj.com/blog/2024/09/03/lsp-good-bad-ugly.html

    Having recently spent a lot more time around typecheckers, I’ve been reading about the Language Server Protocol that glues IDEs and language support services together. The article, from September last year, gives a breakdown of the good and the bad about the protocol, and is a really great dive into the broader topic.

    Much of the pain stems from how the protocol emerged and is managed:

    There is zero open discussion of features before they are added to the spec. Typically they are implemented in VSCode, and then the specification is updated as a fait accompli to document those changes. Implementers of open-source language servers get very influence on the development of the specification.6 There is not even a community space for implementers of language servers to get together and talk about the many tricky corners.

    I feel echoes of this in a lot of different projects I have been around, including PHP internals, ZeroMQ’s protocol, various CNCF working groups, PyTorch and Triton. Protocols and technologies emerge from a need, and grow because that need is shared, but transitioning from a narrow and highly connected problem source to a true standard is difficult: attempt to standardize and bring in voices too early and you just slow down progress to the point something else emerges which solves immediate needs better; leave it too long and the governance questions can be sufficient to encourage folks to rally around forks or alternatives.

    One example of that playing out at the moment is Anthropic’s (and dsp’s!) Model Context Protocol. Tim Kellog wrote a nice post the other day comparing it to OpenAPI, concluding:

    Standards are mostly sociological advancements. Yes, they concern technology, but they govern how society interacts with them. The biggest reason for MCP is simply that everyone else is doing it. Sure, you can be a purist and demand that OpenAPI is adequate, but how many clients support it?

    The reason everyone is agreeing on MCP is because it’s far smaller than OpenAPI. Everything in the tools part of an MCP server is directly isomorphic to something else in OpenAPI. In fact, I can easily generate an MCP server from an openapi.json file, and vice versa. But MCP is far smaller and purpose-focused than OpenAPI is.

  • Scarcity & Abundance

    https://alexdanco.com/2025/03/27/scarcity-and-abundance-in-2025/

    Alex Danco back with a rumination on a what is becoming abundant and that that makes scarce. The two takeaways I had were:

    a) How we think about code changes to a more malleabe, more immediate driver of value rather than some inherent store of value itself.

    b) The easiest place to deploy agents is in areas where you are both already doing the work, but also where that work is executed by somewhat lower-accountability folks (e.g. vendors, contractors, maybe very junior staff).

    It might be a bit too flippant to say, “We’re evolving from a mindset where the codebase is capital (the past few decades of software) and into a mindset where code is labor.” But this is a blog post, so it suits the medium. And it suits today’s energy: new projects and startups are writing a lot more code on the basis of “does this make me money now” (what Simon Wardley would’ve once called “worth-driven development”): the codebase is more like a workforce to be trained than like a factory line to be architected. You still want to put thought in it, but it’s a different kind of thoughtfulness. And you expect profit generation out of it a lot more aggressively than with patient capital deployment.

    The two points are connected. Most very large companies have a real but unspoken delineation between different types of code. Some code is absolutely bullet proof, very closely monitored, and staffed for proactive, ongoing maintenance and enhancement. And some code is… not that.

    We generally don’t explicitly call out the gradations, but we do put different degrees of gating around things. Getting better at both the identification and fencing of critical things (load bearing vs decorative, from a business point of view) will be necessary to unlock the pace of potential improvements via agentic systems, and to focus human attention on the areas where there is a very high inherent complexity/ambiguity.

  • Libraries not Frameworks & Training vs Agentic loops

    A couple of conversations I was in last week around agentic system design reminding me of Brandon Smith’s excellent write libraries, not frameworks post

    A library is a set of building blocks that may share a common theme or work well together, but are largely independent.

    A framework is a context in which someone writes their own code. This could take the form of inversion-of-control, a domain-specific language, or just a very opinionated and internally-coupled library. 

    […]

    So here’s my point: frameworks aren’t always bad, but they are a much bigger risk – for both the creators and the users – than libraries are. If your framework can be a library without losing much, it probably should be

    I feel like we have seen this in ML around general-purpose training loops. Everyone training a model needs a training loop, and there are a lot of commonalities (dataloading, checkpointing, observability and so on). It’s very tempting to build a general training loop that many different groups can use. Unfortunately, this is almost inevitably a framework, rather than a library, and inherently hard to compose.

    When the needs of the modeler exceed the bounds of the framework they either have to make extensive changes or drop the framework and move to a more bespoke set up. In practice this seems to result in a handful of training frameworks that are somewhat domain specialized: for example, a recsys training framework, an LLM training framework, a multimedia/vision oriented framework and so on.

    My sense is the same pattern will happen with agentic loops. A single “one-size-fits-all” agentic framework can feel too broad or rigid, and many teams will carve out domain-specific variants to get the features they truly need. Ideally, we will identify some truly generic components that can be build out library-style, and composed to the domains that we need.

  • A HBS paper on the PyTorch Foundation

    Igniting Innovation: Evidence from PyTorch on Technology Control in Open Collaboration by Daniel Yue, Frank Nagle :: SSRN

    Unexpectedly interesting HBS paper!

    This study looks at the impact of technology control on external contributions in open collaboration contexts by examining the case of PyTorch, a popular machine learning framework, which shifted its governance from a for-profit corporation (Meta) to a non-profit foundation in 2022. The results show that this shift led to a significant decrease in contributions from Meta but a notable increase from external companies.

    The PyTorch project was moved to a foundation in 2022, and that has been a pretty big success (from most any angle you care to look). This paper uses PyTorch as a natural experiment where an already-open-source project changed in governance structure, and what the result was.

    The net result they find is a similar level of overall contributions, but increased contributions from hardware companies. They conclude that there was previously a concern on project direction that could hold up certain types of contributions:

    openness does not magically create incentives for
    external participation without costs, but rather shifts incentives between focal [meaning the originators of the project] and external firms. In particular, control rights theory emphasizes that consideration of the optimal allocation of control rights (with respect to overall welfare) depends on the marginal returns to the ex-ante effort of each party

    There is a lot of adding structure to intuitively reasonable ideas, e.g. that “Users” have a higher incentive to collaborate or contribute as they capture value by by the API, which is unlikely to change regardless of directional shifts by the project owner. “Complementors” on the other hand benefit when their product is used in conjunction with the framework, and therefore they need more ongoing cooperation, so are more sensitive to the control of the project. In PyTorch’s case hardware manufacturers are complementors, and hence their contributions are expected (and did) increase with the changes in governance.

    What’s interesting, as they note in the paper, is that on a technical level the governance of PyTorch hasn’t changed that much. It does seem though a fair conclusion that adding the overall project governance group makes it more likely that such changes could be made, if needed:

    Nevertheless, by changing the governance to a model run by a voting board of other organizations and bringing in the LF, Meta’s singular control of the technical direction of the project (and potentially its social status as the creator of the tool) was greatly diluted.

    My main concern with the conclusions is the confounder of massive extra interest in generative AI post-Chat GPT. They do address that, an attempt to control by looking at TensorFlow:

    usage of AI technologies dramatically increased in December 2022 due to public release of OpenAI’s ChatGPT. And while Chip Manufacturer and Application Developer companies are both affected by this demand shock, it is possible that they are differentially affected by this change in a way that confounds our analysis. To rule out this possibility, we augment our sample by further gathering the external company commits data to TensorFlow, Google’s open source machine learning framework.

    TensorFlow feels like not a great baseline due to the different adoption by the research community. I would be mildly interested to see if XLA had any changes in contributions: the Google diaspora have certainly spread the technology via Jax!

    Regardless, this is an interesting paper and a good contribution in the larger economics of open source. More foundation led projects are better for everyone, so I’m glad to see research in these areas!

  • Vibe Coding & Code Review

    Vibe coding is trendy right now, and it’s part of a bigger shift toward AI-generated code. One of the questions I have about this, at scale, is how it will interact with code review.

    When you go public, you have to comply with Sarbanes-Oxley. SOX mandates a separation of duties, which in practical terms usually means having a process by which changes are tested and approved by someone other than the author before being deployed, and often that is implemented through code review. Companies like Google, Meta, Snapchat, and Uber also have compliance obligations around privacy and security thanks to FTC consent decrees. While these mandates don’t directly say “do code reviews” (they mostly mandate programs) in practice, code reviews are key checkpoints for compliance.

    As we use more model-generated code, traditional code review processes will have to change. Maybe engineers become reviewers of AI-generated code, or we start clearly separating AI-written code from the manually-reviewed stuff—similar to what already happens with some current codegen tools.

    Companies will have to think about ownership and stewardship of changes. Code review requirements like OWNERS or Readability at Google require an approver who isn’t the original author: if the change is LLM authored, is it acceptable for there to be no human author, or is it owned by the person who kicked off the workflow, precluding them from being the reviewer? If we have automated detect-change flows set up (e.g. upgrading downstreams when a dependency changes), is it an individual or a team/oncall that owns the change?

    The bottleneck is human attention and focus. It is plausible to envisage 10-1000x more changes in a large code base, but current review practices simply wouldn’t do an effective job: at best, many changes would be rubber stamped. Work like the diff risk score shows you can do some degree of triage or prioritizing with models. Conceptually you could see extending this to privacy, security and more specific types of risk scoring.

    Work like Policy Zones associates compliance requirements with data and asserts controls in the data flow of the system. This may be easier to scale than trying to validate on code changes.

    Static analysis can also catch patterns of bad usage, like recent work in detecting scraping opportunities. This feels especially important for LLM generated code where a well diffused but bad pattern might be easier to generate than a more novel but safe pattern.

    I expect to see more pressure on developer infrastructure teams to build out capabilities for risk detection, automatic validation and embedding policy information into code or data. There will be an advantage to these being open, industry standard approaches as foundation models will do a better job and require less fine tuning to company specific idioms.

  • Horace He’s thoughts on PyTorch

    https://www.thonking.ai/p/why-pytorch-is-an-amazing-place-to

    I read a the internal version of this (which also doubled as Horace’s leaving note!) and really glad to see he has published it broadly. Its a great look at working in open source and how it tied in to a (very successful) career as an engineer, as well as covering his thoughts on Thinky and the space as a whole!

    In high school, one of the things I feared most was that I would work on some project for 10 years and eventually realize that I’ve wasted my life improving something that nobody cared about. One of the greatest things about working on PyTorch is the certainty that I haven’t.

  • On career growth

    Particularly at the giant tech firms where there are many, many smart people, folks look at success and try to copy it. They set themselves up for promotions against the definition of what is expected at the next level. But those expectations describe an average, not a person. Real people are spiky.

    No one is great at everything, in every situation. Some of the most successful can create situations where they can use their strengths. They shape the work to fit them. Most of us don’t get that luxury, but we can often pick where to play. I’ve taken on projects and roles that looked reasonable, but I knew weren’t a fit, and the results have been from poor-to-fair, never great. That mismatch is costly.

    Choosing projects, teams, or roles is more significant than choosing how to work on then. They decide whether you work with your strengths or against them. Lean into strengths, lean into things that bring you energy and satisfaction. That doesn’t mean staying comfortable. It means knowing where you do your best work and pushing those capabilities or abilities even further, rather than attempting to contort yourself to a generic level N+1.

  • Chinese Tech Term Glossary

    Very interesting list, via Jeff Ding at ChinAI.

    from Chinese primary sources on technology and security, with expert translations and annotations by CSET’s translation team. It’s created and maintained by Ben Murphy with support from the Emerging Technology Observatory.

    https://docs.google.com/spreadsheets/d/15MS8Qp9U-KOaoQF0R_e7lVxF3UBMyZFpC6-cohhnsD0/htmlview#gid=0

  • Engineering Culture at Meta

    A question I’ve been asked recently is how Meta compares to other places I’ve worked, or what makes it different.

    From my conversations and observations, those who disliked working at Meta often cited chaos, short-term focus, and internal politics, while those who liked it called out  autonomy, speed, and the feeling they could work on important projects. To explain that disparity, I refer to three values or themes that shape the work culture.  

    Individuals are responsible for doing impactful work: “impact” is an important concept at Meta, and having a level-appropriate collection of impactful work at performance review time is important for every engineer. If you find yourself in a situation where your impact feels limited, you are generally responsible for exploring ways to address it, or make a change. 

    This means that ICs (individual contributors) at Meta are willing to cross team boundaries to find important work, and will also gravitate towards highly visible projects. They care about how their work is regarded and how it fits in to the wider organization. Internal mobility is fairly easy, so folks will leave teams if they can’t find the right kind of work. Its also reflected in the growth expectations: when I worked at Google and Lyft they also had expectations that IC3s would become IC4s, and IC4s become IC5s (though Google later removed this part), but the timelines were somewhat soft. At Meta, they are firm, and expectations ramp at defined intervals as you approach the boundaries. 

    Practically, this means ICs should expect to identify and collaborate on projects that align with organizational goals and take initiative to push them forward. Managers and leadership provide support, but success heavily depends on individuals ability and desire to chart a path, adapt as needed, and ensure their contributions are visible. In general, there’s a strong bias towards getting things done, getting things out,  “rough consensus and running code”.

    Dave Anderson has written about how much more helpful he found teams at Meta, vs Amazon where there was a lot of horse trading for collaboration. Part of that is driven, I think, by this responsibility for impact. Having another team’s thanks, or enabling results for them, allows you to claim some credit for their impact with relatively little effort. Conversely, intentionally blocking another team can be seen as gatekeeping, which is frowned upon.

    No gatekeeping: Some version of “Move Fast” has been in Meta’s official values for a long time, and the company still operates at a good clip. Part of that is aided by generally making it easy to go make changes wherever they need to be made. One example of this I use with Google folks is OWNERS files. Google and Meta are both monorepo based, but at Google you have sets of services with clear owners, and touching code in another team’s service requires their full blessing. Meta also operates a monorepo, but there is much more fluidity — folks can land changes anywhere they need to. In part this is because the original Meta product, Facebook itself, is a monolith, but there is a deeper cultural aspect here.

    For example, even very senior engineers will very rarely say “no” to something. Instead of outright rejection, feedback is often framed as suggestions or concerns to consider. This encourages risk-taking and innovation, but it also places a significant responsibility on engineers to weigh feedback carefully, address risks , and seek out champions from stakeholders. This is one source of the disconnect between folks who see Meta as a place where it’s ok to take risks and folks who don’t: if you take a risk, it fails, and at PSC (performance review) time someone affirms they called out the problems that occurred and you didn’t take appropriate measures, you will be dinged. I have seen people interpret this kind of feedback as nits or suggestions, rather than weighing it heavily and convincing others that the risk is well managed ahead of time.

    In general, folks are expected to be helpful, to provide guidance to others, and not to put up walls, so it can be a tricky balance when outside teams or other engineers come in and ride roughshod over a team’s plans or projects. As an engineer, escalating misalignments on goals/priorities to management is usually well supported, and as a manager, putting engineers together to get to technical solutions across teams is expected. Enacting hard blocks where one team can’t achieve their goals because another was in the way is less so.

    The heavy dependence on individuals and relationships, particularly for cross-team projects, is another key theme:

    It’s a social company: Somewhat unsurprisingly for a company that started around a social network, Meta is a pretty social company. The internal social network is a firehose of information, and there are deep networks of connections across the company between senior engineers, managers and executives.

    Its important for engineers to talk about their work in order to find folks interested in it, build connections and relationships with them, and have a good sense not just of their org but the universe of organizations that they operate within. A fairly common failure pattern is to build a good relationship with one side of an org and ignore another, developing a sizable blind spot that later comes back to be a problem.

    The official org chart is the secondary and lagging structure at the company. The more important structure is the informal network of relationships that where many things get done, and decisions get made.

    This dynamic can sometimes feel political, which is why some describe Meta this way. While there are large-company politics (it is a large, influential company), for most people, it’s less about traditional power politics and more about navigating cliques and informal networks of folks who have worked together on multiple projects and have mutual trust and respect. The company leans a lot on strong, senior engineers to drive projects to success, and those folks may not report into the org that is officially doing the work, or may be at a lower or higher level in the org chart than you might expect. They will often work by going directly to the people they know to unblock issues, drive important changes, or get alignment on a controversial decisions.

    For big enough changes, org structural changes do follow, but they usually lag rather than lead the work itself.

    What are the downsides and upsides?

    As folks who have had a bad time can attest, Meta can be chaotic. There can be parallel implementations, people can swarm on important projects to the detriment of those trying to work on them, and less impactful projects can end up unowned and passed around. It can feel like information overload, with more being published than you can possibly follow. At the same time, there’s can be an information drought when truly important conversations happen in small, exclusive groups. For example, I’ve seen feedback about a project be shared openly between a small collection of engineers and leaders, without ever clearly reaching the team responsible for developing it.

    The flip side is that this combination of values allows Meta to pivot surprisingly fast. Changes that would have required months at some companies can be kicked off in a day, particularly by very senior, well-connected leaders. Senior ICs can take problem descriptions, and quickly form an idea of who might have thoughts on it, and pull them in. Soon you have a loose group (often later structured into a “v-team” or virtual team) that can quickly align and drive change. The lack of gatekeeping, both technically and culturally, reduces the corporate immune reaction to large changes. The incentive to individual impact encourages folks to jump onboard important things without having to work out if that means a team change, or what their long-term situation might be.

  • Daniel Schmachtenberger’s Questions

    Daniel’s list of questions

    Questions that can support better understanding, to inform better strategy and design.

    Daniel Schmachtenberger speaks compellingly on the value systems implied and expressed by any piece of technology. He has a list of questions to help think more broadly around something when introducing it.

  • Timeline on the Google Reader shutdown

    Interesting to read this breakdown of what was happening with Reader: I recall the discussions internally at Google and knew some folks floating around it (including at least one person who volunteered to take on maintenance) but this was largely new to me. Uncomfortably enough was on stage giving a talk for Google the day the shut down happened. I asked for questions, got a lot of hands, then qualified I was looking for questions about the talk and not Reader, and most of the hands went down.

    https://blog.persistent.info/2013/06/google-reader-shutdown-tidbits.html?m=1